Senior Site Reliability Engineer (Bengaluru)

Senior Site Reliability Engineer (Bengaluru)

03 Sep
|
Epsilon Data Management
|
Bengaluru

03 Sep

Epsilon Data Management

Bengaluru

Job Summary

Overview: Epsilon is seeking a Senior Site Reliability Engineer to help build, operate, and evolve highly scalable, resilient, and secure cloud platforms supporting critical enterprise applications. As part of a large-scale cloud transformation initiative, you will partner closely with Engineering, DevOps, Platform, and Security teams to establish reliability practices, improve operational excellence, and ensure systems meet performance, availability, and scalability objectives. This is a hands-on technical leadership role requiring deep expertise in cloud infrastructure, Kubernetes, observability, incident management, and reliability engineering.

You will drive technical decisions, influence engineering practices, and help teams design systems that are resilient by design.

Responsibilities

- Design and implement reliability strategies for distributed systems running across AWS and GCP.

- Define and measure Service Level Indicators (SLIs), Service Level Objectives (SLOs), and reliability metrics.

- Build and enhance observability solutions using monitoring, logging, tracing, and alerting platforms.

- Lead incident response, root cause analysis, and postmortem processes to improve system reliability.

- Collaborate with engineering teams to improve system performance, resiliency, scalability, and operational readiness.

- Automate operational processes and reduce toil through engineering solutions.

- Guide teams on reliability-focused architecture decisions, capacity planning, and non-functional requirements.

- Experience with configuration management and automation tools (e.g., Ansible, Puppet, Chef).

- Strong experience with Linux system administration including user management, file systems, networking, and performance tuning.

- Experience managing containers and orchestration (Docker, Kubernetes) is a plus.

- Familiarity with monitoring and logging tools (Nagios, Prometheus, Grafana, ELK stack).

- Experience with databases such as MySQL, PostgreSQL, or equivalent tools.

- Develop and enforce system standards, automation practices,



and recommend improvements to enhance performance and reliability.

- Install, configure, and maintain commercial and open-source applications on Linux operating systems (e.g., RHEL, CentOS, Ubuntu).

- Manage and support BAU (Business As Usual) operational activities.

- Handle ServiceNow incidents, requests, and change tickets.

- Perform Linux operating system fixing and issue resolution.

- Conduct root cause analysis (RCA) for incidents and service disruptions.

- Ensure infrastructure stability, reliability, and service availability.

- Drive automation initiatives to improve operational efficiency and reduce manual effort.

- Collaborate with multi-functional teams to support enterprise infrastructure and platform services.

- Monitor system health, performance, and capacity using observability and monitoring tools.

- Implement OS patching activities across Linux servers and hybrid infrastructure environments.

- Drive vulnerability remediation efforts to address security findings and compliance requirements.

- Coordinate with application owners and partners for patch validation and deployment activities.

- Ensure consistency to security, compliance, and operational standards.

- Automate routine operational tasks using Python, Ansible, Terraform, Bitbucket, PowerShell, ELK Stack, Shell / Bash Scripting.

- Monitor system health, performance, availability, and capacity using enterprise monitoring and observability platforms.

- Support cloud and on-premises infrastructure across AWS, Azure, GCP, VMware, and enterprise data centre environments.

- Drive continuous service improvements through automation, standardization, and proactive problem management.





- Manage the complete ELK (Elasticsearch, Logstash, Kibana) platform lifecycle, including deployment, upgrades, maintenance, capacity planning, and decommissioning.

- Design, develop, and maintain Kibana dashboards, visualizations, alerts, and reporting solutions for infrastructure, security, and operational monitoring.

- Ensure log ingestion, retention, performance optimization, data availability, and platform reliability across the ELK ecosystem.

- Fix and resolve issues related to Elasticsearch clusters, Logstash pipelines, and Kibana dashboards.

- Hands-on experience with Dell and HPE server hardware troubleshooting, health checks, firmware management, and diagnostics using iDRAC, iLO, and vendor management tools.

Qualifications

- Education, Experience, and Licensing Requirements Bachelor s degree in computer science or related field (or equivalent experience).

- Minimum of 5+ years of proven experience in Linux system administration or IT infrastructure.

- Experience working across multiple operating systems with robust emphasis on Linux platforms.

- Relevant certifications preferred: Red Hat Certified System Administrator (RHCSA) or Engineer (RHCE) Linux+ or equivalent VMware or cloud certifications (AWS, Azure) are a plus.

- Python, Ansible, Terraform, Bitbucket, PowerShell, ELK Stack, Shell Scripting (Bash, Python comparable languages).

- Set Yourself Apart With Experience supporting large-scale cloud migration or modernization programs.

- Expertise in incident management and production operations for high-availability systems.

- Experience implementing chaos engineering or resilience testing practices.

- RHEL/AWS/AZURE/GCP certifications.

- Experience working in Agile, DevOps, or DevSecOps environments.

Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.

📌 Senior Site Reliability Engineer (Bengaluru)
🏢 Epsilon Data Management
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior site reliability engineer (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: senior site reliability engineer (bengaluru) / bengaluru