Site Reliability Engineer - Information Technology (Chennai)

Site Reliability Engineer - Information Technology (Chennai)

22 Sep
|
Apollo Hospitals
|
Chennai

22 Sep

Apollo Hospitals

Chennai

The Apollo Hospitals Family

Set-up in 1983 by Dr. Prathap C. Reddy, renowned as the architect of modern healthcare in India. As the nations first corporate hospital, Apollo Hospitals is acclaimed for pioneering the private healthcare revolution in the country.

Apollo Hospitals has emerged as Asia’s foremost integrated healthcare services provider and has a robust presence across the healthcare ecosystem, including Hospitals, Pharmacies, Primary Care & Diagnostic Clinics and several retail health models. The cornerstones of Apollo’s legacy are its unstinting focus on clinical excellence, affordable costs, modern technology and forward-looking research & academics.

Role Description

The Site Reliability Engineer is responsible for proactively monitoring, diagnosing and resolving production issues across Apollo's clinical, revenue and enterprise applications running across 50+ hospitals. Operating in a three-shift plus reliever model to provide 24x7 coverage, the SRE ensures that defined SLIs/SLOs are met, error budgets are governed, and permanent fixes are driven upstream into engineering. The role requires balancing reactive incident response with proactive reliability engineering — automating runbooks, tuning observability, and participating in chaos and disaster recovery drills.

Detailed Job Responsibility & Accountabilities

- Operate production monitoring consoles (Dynatrace, Azure Monitor, Log Analytics) during assigned shift and proactively detect anomalies before user impact.
- Triage and respond to Sev-1, Sev-2 and Sev-3 incidents per defined SLAs; coordinate with application,



infrastructure, network, database and vendor teams for resolution.
- Execute documented runbooks for common failure modes and continuously improve/automate them.
- Participate in incident command during major outages; own communication updates to clinical and business stakeholders during critical events.
- Perform blameless postmortems for Sev-1/Sev-2 incidents and drive corrective actions to closure.
- Instrument new services with golden signals, custom metrics, distributed tracing and meaningful alerting.
- Identify and eliminate toil through scripting, automation and self-service tooling.
- Participate in chaos engineering, game-days and DR drills for clinical platforms.
- Maintain on-call readiness, escalation paths and shift handover quality; publish shift summary reports.
- Contribute to reliability scorecards, error budget tracking and availability reports for leadership.
- Support release readiness reviews and post-deployment verification for high-risk changes.

Desired Skills & Competencies

- Must have 4+ years of experience in SRE, production support engineering or DevOps with incident response exposure.
- Must have hands-on experience with observability tools such as Dynatrace,



Prometheus/Grafana, ELK, Splunk, Azure Monitor.
- Must have robust Linux/Windows troubleshooting, networking fundamentals, and database (SQL, PostgreSQL, MongoDB) query skills.
- Must have scripting proficiency in Python, PowerShell, Bash.
- Must have working knowledge of Kubernetes, Docker, Azure cloud services.
- Must have exposure to ticketing (ServiceNow, Jira), alerting (PagerDuty, Opsgenie) and communication discipline during incidents.
- Preferred exposure to SLO engineering and error budget practice.
- Comfort with rotating shift schedules (3-shift + reliever model) is essential.

Job Performance Parameters

- SLA adherence for incident acknowledgement, triage and resolution
- % incidents detected proactively before user impact during own shift
- MTTR contribution and number of permanent fixes raised to engineering
- Toil reduction through automation (hours saved/quarter)
- Shift handover quality and runbook compliance
- Postmortem action closure rate

Managerial and Behavioral Competencies

- Calm under pressure and structured problem solving
- Effective Communication during high-severity incidents
- Team Spirit and shift collaboration
- Ownership and follow-through
- Continuous learning mindset

Preferred Language / Industry / Location / Desired Work Experience

English, Tamil, Hindi / Healthcare + IT / Chennai / 4+ years

Business Unit, Work Location & Work Timing

Corporate / Chennai - Ali Towers / Rotational shifts in a 3-shift plus reliever model providing 24x7x365 coverage

📌 Site Reliability Engineer - Information Technology (Chennai)
🏢 Apollo Hospitals
📍 Chennai

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer - information technology (chennai) / chennai