24 Sep
|
The Hartford
|
Secunderabad
24 Sep
The Hartford
Secunderabad
About the Role
We are looking for a Lead Reliability Engineer to own and evolve The Hartford's enterprise Software Delivery Framework (SDF) platform. You will be the technical authority on Jenkins (primary), Harness, Git / GitHub Enterprise, uDeploy, Nexus, and SonarQube — driving platform stability, service reliability, and cost-efficiency through robust operational engineering and governance. An AI-first mentality is a core expectation. Cross-team collaboration is a must — this role partners with engineering, security, observability, release management, and infrastructure teams to align standards and unblock delivery.
Key Responsibilities
- Own the full SDF platform lifecycle: Jenkins, Harness, Git / GitHub Enterprise, uDeploy, Nexus, and SonarQube.
- Ensure platform stability and availability across SDF tooling through proactive reliability engineering practices.
- Define and enforce enterprise reliability standards, operational controls, and compliance guardrails.
- Drive incident prevention and rapid recovery: observability, early risk detection, runbook maturity, and resilience testing.
- Apply AI-first approaches for reliability diagnostics, intelligent alert triage, and auto-remediation opportunities.
- Own end-to-end RCA for Sev1/Sev2 SDF incidents — from detection through corrective action and verified closure; publish stakeholder summaries.
- Drive alert noise reduction across SDF tooling — enforce signal-to-noise standards and measurable alert quality gates.
- Lead cost optimization initiatives across tooling and infrastructure without compromising reliability or developer experience.
- Coach junior engineers — define escalation criteria, build triage playbooks, and close knowledge gaps.
- Cross-team collaboration — partner with application engineering, observability, security, release management, and infrastructure teams to align standards and unblock delivery.
- Lead end-user experience outcomes as the centralized SDF platform owner, ensuring every reliability and stability improvement delivers a simpler, faster, and more consistent experience across all SDF tools.
- Build and maintain a self-onboarding framework so new teams can independently adopt SDF operational best practices.
- Manage SDF infrastructure via Terraform and Ansible; automate provisioning, health checks, and self-healing.
- Track and report platform reliability and efficiency metrics: availability, MTTR, incident recurrence, and cost savings.
Platform & Tools
Jenkins Harness Git / GitHub Enterprise uDeploy Nexus SonarQube Terraform Ansible Harness Python / Groovy / Bash REST APIs Rally
Requirements
- Jenkins administration (required)
- Harness CD pipelines
- Git / GitHub Enterprise
- uDeploy
- Nexus Repository Manager
- SonarQube
- Platform reliability engineering and operational governance
- Service stability, resilience, and observability practices
- Cross-team collaboration (required)
- End-to-end RCA ownership
- Cost optimization and efficiency mindset
- Coaching and mentoring junior engineers
- Terraform / Ansible (IaC)
- AI-first problem-solving approach
- Regulated industry experience a plus
📌 Lead Reliability Engineer (Secunderabad)
🏢 The Hartford
📍 Secunderabad