Overview
About Business Unit:
At the core of all that Epsilon does is a team that sets the foundation of our IT infrastructure. The team drives innovation and efficiency through pioneering technology across Epsilon's platforms and business verticals. From being the first point of contact for infrastructure needs to final deployment, the team provides end-to-end solutions for our client-facing platforms. ETS supports all aspects of revenue-generating platforms for Epsilon and sets the architectural direction for our enterprise deployments. By adopting the newest technologies, such as Cloud, Automation, and Artificial Intelligence, the team is at the front of redefining our digital business and capturing current opportunities.
Overview:
Epsilon is seeking a
Senior Site Reliability Engineer to help build, operate, and evolve highly scalable, resilient, and secure cloud platforms supporting critical enterprise applications. As part of a large-scale cloud transformation initiative, you will partner closely with Engineering, DevOps, Platform, and Security teams to establish reliability practices, improve operational excellence, and ensure systems meet performance, availability, and scalability objectives.
This is a hands-on technical leadership role requiring deep expertise in cloud infrastructure, Kubernetes, observability, incident management, and reliability engineering. You will drive technical decisions, influence engineering practices, and help teams design systems that are resilient by design.
Job Description: - Senior Site Reliability Engineer
- 5+ years of proven experience in Site Reliability Engineering, Cloud Engineering, DevOps, or Platform Engineering.
- 5+ years of proven experience supporting enterprise Linux environments (RHEL, CentOS, Alma) across multiple sites and centralized services.
- 3+ years of proven experience implementing virtualization solutions using
VMware vSphere, KVM, Proxmox or similar technologies with high availability features (vMotion, clustering, failover).
- 3+ years of proven experience in configuring and setting up workflow automation (e.g., N8N).
- Experience with configuration management and automation tools (e.g., Ansible, Puppet, Chef).
- Advanced hands-on engineering with senior ownership of infrastructure design, automation, and service delivery.
- Core Tools:Python, Ansible, Terraform, Bitbucket, PowerShell, ELK Stack, Shell Scripting (Bash, Python),Enterprise Monitoring & Observability Platforms,n8n
- Strong experience with
Linux system administration including user management, file systems, networking, and performance tuning.
- Experience managing
containers and orchestration (Docker, Kubernetes) is a plus.
- Familiarity with
monitoring and logging tools (Kibana, Grafana, ELK stack).
- Demonstrated experience using ITIL-based ticketing systems (e.g., JIRA, ServiceNow).
- Strong experience supporting production systems in AWS/GCP/Azure environments.
- Deep understanding of SRE principles, including SLIs, SLOs, error budgets, and operational excellence.
- Experience operating and fixing issues on Kubernetes platforms such as EKS and/or GKE.
- Strong knowledge of observability tools such as, Grafana, CloudWatch, Cloud Monitoring, Datadog, Splunk, or similar.
- Experience with Infrastructure as Code tools such as Terraform.
- Strong scripting and automation skills using Python, Ansible, Terraform, Bitbucket, PowerShell, ELK Stack, Shell Scripting (Bash, Python comparable languages.
- Core Tools:Python, Ansible, Terraform, Bitbucket, PowerShell, ELK Stack, Shell Scripting (Bash, Python),Enterprise Monitoring & Observability Platforms,n8n
- Solid understanding of networking, distributed systems, cloud security, and performance optimization.
- Primary Expertise:Linux Server Engineering, Automation, DevOps, Monitoring, and Observability
- Linux Platforms:Linux (RHEL, CentOS, AlmaLinux)
- Technical Depth:Advanced hands-on engineering with senior ownership of infrastructure design, automation, and service delivery.
Click here to view how Epsilon transforms marketing with 1 View, 1 Vision and 1 Voice.
Responsibilities
- Design and implement reliability strategies for distributed systems running across AWS and GCP.
- Define and measure Service Level Indicators (SLIs), Service Level Objectives (SLOs), and reliability metrics.
- Build and enhance observability solutions using monitoring, logging, tracing, and alerting platforms.
- Lead incident response, root cause analysis, and postmortem processes to improve system reliability.
- Collaborate with engineering teams to improve system performance, resiliency, scalability, and operational readiness.
- Automate operational processes and reduce toil through engineering solutions.
- Guide teams on reliability-focused architecture decisions, capacity planning, and non-functional requirements.
- Experience with configuration management and automation tools (e.g., Ansible, Puppet, Chef).
- Strong experience with Linux system administration including user management, file systems, networking, and performance tuning.
- Experience managing containers and orchestration (Docker, Kubernetes) is a plus.
- Familiarity with monitoring and logging tools (Nagios, Prometheus, Grafana, ELK stack).
- Experience with databases such as MySQL, PostgreSQL, or equivalent tools.
- Develop and enforce system standards, automation practices, and recommend improvements to enhance performance and reliability.
- Install, configure, and maintain commercial and open-source applications on Linux operating systems (e.g., RHEL, CentOS, Ubuntu).
- Manage and support BAU (Business As Usual) operational activities.
- Handle ServiceNow incidents, requests, and change tickets.
- Perform Linux operating system fixing and issue resolution.
- Conduct root cause analysis (RCA) for incidents and service disruptions.
- Ensure infrastructure stability, reliability, and service availability.
- Drive automation initiatives to improve operational efficiency and reduce manual effort.
- Collaborate with multi-functional teams to support enterprise infrastructure and platform services.
- Monitor system health, performance, and capacity using observability and monitoring tools.
- Implement OS patching activities across Linux servers and hybrid infrastructure environments.
- Drive vulnerability remediation efforts to address security findings and compliance requirements.
- Coordinate with application owners and partners for patch validation and deployment activities.
- Ensure consistency to security, compliance, and operational standards.
- Automate routine operational tasks using , Python, Ansible, Terraform, Bitbucket, PowerShell, ELK Stack, Shell / Bash Scripting.
- Monitor system health, performance, availability, and capacity using enterprise monitoring and observability platforms.
- Support cloud and on-premises infrastructure across AWS, Azure, GCP, VMware, and enterprise data centre environments.
- Drive continuous service improvements through automation, standardization, and proactive problem management
- Manage the complete ELK (Elasticsearch, Logstash, Kibana) platform lifecycle, including deployment, upgrades, maintenance, capacity planning, and decommissioning.
- Design, develop, and maintain Kibana dashboards, visualizations, alerts, and reporting solutions for infrastructure, security, and operational monitoring.
- Ensure log ingestion, retention, performance optimization,
data availability, and platform reliability across the ELK ecosystem.
- Fix and resolve issues related to Elasticsearch clusters, Logstash pipelines, and Kibana dashboards.
- Hands-on experience with Dell and HPE server hardware troubleshooting, health checks, firmware management, and diagnostics using iDRAC, iLO, and vendor management tools.
Qualifications
Education, Experience, and Licensing Requirements
- Bachelor's degree in computer science or related field (or equivalent experience).
- Minimum of
5+ years of proven experience in Linux system administration or IT infrastructure.
- Experience with databases such as MySQL, PostgreSQL, or equivalent tools.
- Experience working across multiple operating systems with strong emphasis on Linux platforms.
- Relevant certifications preferred:
- - Red Hat Certified System Administrator (
RHCSA) or Engineer (
RHCE)
- Linux+ or equivalent
- VMware or cloud certifications (AWS, Azure) are a plus
- Python, Ansible, Terraform, Bitbucket, PowerShell, ELK Stack, Shell Scripting (Bash, Python comparable languages.
Set Yourself Apart With
- Experience supporting large-scale cloud migration or modernization programs.
- Expertise in incident management and production operations for high-availability systems.
- Experience implementing chaos engineering or resilience testing practices.
- RHEL/AWS/AZURE/GCP certifications.
- Experience working in Agile, DevOps, or DevSecOps environments.
Additional Information
Epsilon is a global data, technology and services company that powers the marketing and advertising ecosystem. For decades, we've provided marketers from the world's leading brands the data, technology and services they need to engage consumers with 1 View, 1 Vision and 1 Voice. 1 View of their universe of potential buyers. 1 Vision for engaging each individual. And 1 Voice to harmonize engagement across paid, owned and earned channels.
Epsilon's comprehensive portfolio of capabilities across our suite of digital media, messaging and loyalty solutions bridge the divide between marketing and advertising technology. We process 400+ billion consumer actions each day using advanced AI and hold many patents of proprietary technology, including real-time modeling languages and consumer privacy advancements. Thanks to the work of every employee, Epsilon has been consistently recognized as industry-leading by Forrester, Adweek and the MRC. Epsilon is a global company with more than 9,000 employees around the world.
Our pillars aren't just words. They're how we show up every day.
- People centricity: We focus on employee well-being in an setting where colleagues truly care about each other.
- Collaboration: We work together, support one another, and collectively achieve goals.
- Growth: There are endless opportunities for growth through learning, development and career advancement.
- Innovation: We drive progress through cutting-edge solutions and forward-thinking approaches.
- Flexibility: We've created a balance between work and personal life, and we encourage adaptability to solve problems creatively.
Our values guide us to create value for our clients, our people and consumers.
- Act with integrity
- Work together to win together
- Innovate with purpose
- Respect all voices
- Empower with accountability
These pillars and values are our foundation-shaping our culture, guiding our decisions, and uniting us in common purpose.
Epsilon is an Equal Opportunity Employer.
Epsilon is committed to promoting diversity, inclusion, and equal employment opportunities by using reasonable efforts to attract, recruit, engage and retain qualified individuals of all ethnicities and backgrounds, including, but not limited to, women, people of color, LGBTQ individuals, people with disabilities and any other underrepresented groups, traits or characteristics.
📌 Senior Site Reliability Engineer (India)
🏢 Epsilon
📍 India