We are looking for an experienced SRE Manager to lead our Site Reliability Engineering team. The ideal candidate will have a robust background in DevOps practices, system reliability, and team leadership.
Key Responsibilities
Lead, mentor, and manage a team of SRE/DevOps engineers.
Define and implement SRE best practices (SLIs, SLOs, error budgets).
Ensure system reliability, scalability, and performance.
Drive automation initiatives.
Collaborate with cross-functional teams.
Own CI/CD pipelines and release management.
Lead incident response and RCA processes.
Establish monitoring and observability frameworks.
Manage cloud infrastructure (AWS/Azure/GCP).
Implement disaster recovery plans.
Required Skills & Qualifications
7+ years of experience in SRE/DevOps roles.
3+ years of team management experience.
Experience with cloud platforms (AWS/Azure/GCP).
Knowledge of CI/CD tools (Jenkins, GitLab CI).
Experience with Docker and Kubernetes.
Scripting skills (Python, Bash).
Knowledge of Monitoring tools (Prometheus, Grafana, ELK).
Preferred Qualifications
Experience with microservices.
Cloud certifications are a plus.
Solid problem-solving skills.