Senior Service Reliability Engineer Job Description
Key Responsibilities
Application Reliability, Availability & Security
- Own end-to-end reliability of applications in production environments
- Ensure adherence to SLA, SLO, and SLI targets
- Continuously improve availability, latency, and performance metrics
- Ensure applications comply with security, privacy, and governance standards
- Support patching, vulnerability remediation, and compliance requirements
Incident Management & Problem Resolution
- Lead incident response, triage, and resolution across application layers
- Perform root cause analysis (RCA) and drive permanent fixes
- Reduce Mean Time to Recovery (MTTR) through automation and process improvements
- Act as escalation point for critical production issues
Automation & Reliability Engineering
- Develop automation for deployment, monitoring, and recovery processes
- Drive “Reliability as Code” and infrastructure automation
- Build self-healing mechanisms and reduce manual operational effort
- Design and maintain CI/CD pipelines for application delivery
- Ensure reliable and consistent deployments using automated pipelines
- Support application release cycles with zero/low downtime strategies
Observability & Capacity Planning
- Implement monitoring, logging, and alerting systems
- Define meaningful alerts and reduce noise/false positives
- Create dashboards and metrics for real-time health visibility
- Conduct performance testing and tuning
- Forecast capacity and ensure scalability of applications
- Optimize cost vs performance in cloud environments
Required Qualifications
Education
- Bachelor’s/Master’s degree in Computer Science or related field
Experience
- 8+ Years in SRE / DevOps / Production Engineering roles
- Hands-on experience supporting production-grade applications in cloud environments