We are seeking a talented individual to join our GIS at Marsh. This role will be based in Mumbai. This is a hybrid role that has a requirement of working at least three days a week in the office.
Senior Manager – Site Reliability Engineering (SRE) / Reliability Operations
Role Overview
We are looking for a Senior Manager – SRE / Reliability Operations to drive reliability engineering practices across critical services and platforms. In this role, you will establish and mature SLI/SLO/SLA and error budget governance, lead operational excellence (ITIL-aligned incident/problem/change), and partner closely with engineering, architecture, and DevOps teams to reduce toil through automation and improve resilience, observability, and release reliability.
We will count on you to:
- SRE fitness: SLI/SLO/SLA & error budget management
- Define, implement, and continuously improve SLIs and SLOs for services/products in collaboration with product and engineering teams
- Align SLAs with business expectations and ensure operational commitments are measurable and reportable
- Create and manage error budgets, including:
- Error budget policies and burn-rate thresholds
- Release gating recommendations based on error budget status
- Executive reporting on reliability posture and trade-offs
- Build SLO dashboards and reliability scorecards for leadership and stakeholders
- Conduct periodic SLO reviews and drive corrective actions when reliability trends degrade
- Software engineering, automation, and toil reduction
- Identify operational toil and repetitive manual work; design and deliver automation to reduce or eliminate it
- Develop tooling/scripts/services using standard languages (e.g., Python, Go, Java, PowerShell, or similar) aligned to the team’s stack
- Implement self-healing patterns (automated remediation, auto-rollback, auto-scaling, secure retries)
- Standardize operational runbooks and embed automation into runbooks wherever possibl