19 Sep
|
Eli Lilly
|
Hyderabad
19 Sep
Eli Lilly
Hyderabad
Role summary
You build the self-healing automation, author the runbooks, and turn root-cause analysis into durable engineering fixes that let the production estate heal itself instead of paging a human. You work within the standards the Senior Principal SRE Engineer sets - SLOs, error budgets, observability - and youre the one who encodes them into working automation and documented procedure.
This is a hands-on individual-contributor role focused on execution and codification rather than cross-estate reliability strategy. You decide, in partnership with the Senior Principal SRE Engineer, which recurring patterns warrant a self-healing investment versus a documented manual runbook, and you build whichever is right.
You are an individual contributor. You do not manage people. You partner daily with the Senior Principal SRE Engineer, the agentic automation engineering team, and Operations on validating outcomes. Success is measured by self-healing coverage, runbooks authored and adopted, reduced recurrence of known failure modes, and the safety record of every automation you sign off.
What youll be doing
Self-healing automation & resilience patterns
- Design and build self-healing automation - circuit breakers, graceful degradation, automated remediation - for the failure modes that recur most across the estate.
- Run resilience or chaos testing to validate that self-healing patterns behave correctly before theyre trusted in production.
- Continuously expand self-healing coverage as new failure modes are identified and proven safe to automate.
- Partner with the Senior Principal SRE Engineer on which failure modes justify self-healing investment versus a documented manual runbook.
Runbook authorship & validation
- Author and validate the remediation runbooks for the production estate: safe execution order, rollback steps, and exception handling for every documented fix.
- Keep the runbook library current as systems, dependencies, and failure modes evolve, retiring runbooks that no longer apply.
- Define and apply the graduation criteria that let a runbook move from human-executed to agent-assisted to autonomous.
RCA to durable fix
- Lead or contribute to root-cause analysis for significant incidents,
and drive the blameless postmortem process to a durable engineering fix - not just a narrative.
- Convert recurring incident patterns into codified runbooks and, where appropriate, self-healing automation.
- Track fix effectiveness against recurrence, and escalate to the Senior Principal SRE Engineer when a fix needs broader engineering investment.
- Participate in high-severity incident response, including acting as incident commander for escalations within your area.
Partnership with agentic automation & operations
- Partner with the Agentic Automation Engineering team on which fixes are safe to hand off as agent-assisted remediations, and on the confidence thresholds and human-in-the-loop boundaries that keep them safe.
- Sign off on agent graduation criteria (accuracy over volume, zero P1/P2 caused) before an automation moves to a higher autonomy tier.
- Partner with Operations on outcome validation, feeding whats learned back into the runbook library and self-healing patterns.
Incident response & regulated-environment practice
- Ensure runbooks and self-healing automation meet Lillys change-control, audit, and validated-environment standards.
- Document procedures so that audit evidence falls out of normal operation, not a special exercise.
- Mentor other reliability and automation engineers on runbook quality and self-healing design.
- Contribute proven patterns back to the broader reliability practice, in partnership with the Senior Principal SRE Engineer and Senior Architect.
How you will succeed
At the principal engineering level for reliability, success is defined by the durability and safety of what you build:
- Be recognized as the engineer who turns incidents into durable fixes, not repeat pages.
- Demonstrate measurable growth in self-healing coverage and runbook adoption, with falling recurrence of known failure modes.
- Maintain a clean safety record: automations you sign off dont cause P1/P2 incidents.
- Build runbooks and automation that make good practice the default, not a personal habit.
What you should bring
Required
- 10+ years of progressive engineering experience, with meaningful time as a Site Reliability Engineer, Production Engineer, or equivalent, including hands-on ownership of self-healing automation or runbook-driven remediation for a multi-application production estate.
- Production reliability experience in a regulated or audited environment (GxP, SOX, HIPAA, PCI, or equivalent), including familiarity with change-control discipline, audit evidence, and validated-system constraints.
- Hands-on experience authoring and validating runbooks: safe execution order, rollback steps, and exception handling for real remediation procedures.
- Deep, hands-on fluency across the SRE technical stack: observability (Splunk, Datadog, Recent Relic, or Grafana/Prometheus); infrastructure-as-code (Terraform); CI/CD pipeline hardening; Kubernetes and container platforms; and at least one major public cloud (AWS, Azure, or GCP).
- Practical experience designing self-healing patterns (circuit breakers, graceful degradation, automated remediation) and validating them before theyre trusted in production.
- Track record of contributing meaningfully to root-cause analysis and blameless postmortems, converting findings into durable engineering fixes measured by reduced recurrence rather than narrative quality.
- Bachelors degree or higher in Computer Science, Information Technology, or a closely related engineering field.
Preferred
- Hands-on experience designing self-healing automation and running chaos engineering or resilience-testing programs (AWS Fault Injection Service, Gremlin, LitmusChaos, or equivalent) tied to measurable reliability gains.
- Deep AWS fluency across reliability-relevant services (EKS, ECS, Lambda, CloudWatch, X-Ray, Systems Manager, Route 53), and familiarity with AWS Well-Architected Reliability Pillar.
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Principal SRE Engineer (Hyderabad)
🏢 Eli Lilly
📍 Hyderabad