Experience : 5+ years (with 3+ years in team management).
Location : Gurgaon (WFO).
Role Overview
We are looking for an experienced SRE Manager to lead our Site Reliability Engineering team. The ideal candidate will have a strong background in DevOps practices, system reliability, and team leadership.
Key Responsibilities
- Lead, mentor, and manage a team of SRE/DevOps engineers.
- Define and implement SRE best practices (SLIs, SLOs, error budgets).
- Ensure system reliability, scalability, and performance.
- Drive automation initiatives.
- Collaborate with cross-functional teams.
- Own CI/CD pipelines and release management.
- Lead incident response and RCA processes.
- Establish monitoring and observability frameworks.
- Manage cloud infrastructure (AWS/Azure/GCP).
- Implement disaster recovery plans.
Required Skills & Qualifications
- 7+ years of experience in SRE/DevOps roles.
- 3+ years of team management experience.
- Experience with cloud platforms (AWS/Azure/GCP).
- Knowledge of CI/CD tools (Jenkins, GitLab CI).
- Experience with Docker and Kubernetes.
- Scripting skills (Python, Bash).
- Knowledge of Terraform/CloudFormation.
- Monitoring tools (Prometheus, Grafana, ELK).
Preferred Qualifications
- Experience with microservices.
- Cloud certifications are a plus.
- Robust problem-solving skills.