We are looking for an experienced SRE Manager to lead our Site Reliability Engineering team. The ideal candidate will have a robust background in DevOps practices, system reliability, and team leadership.
Key Responsibilities
• Lead, mentor, and manage a team of SRE/DevOps engineers.
• Define and implement SRE best practices (SLIs, SLOs, error budgets).
• Ensure system reliability, scalability, and performance.
• Drive automation initiatives.
• Collaborate with cross-functional teams.
• Own CI/CD pipelines and release management.
• Lead incident response and RCA processes.
• Establish monitoring and observability frameworks.
• Manage cloud infrastructure (AWS/Azure/GCP).
• Implement disaster recovery plans.
Required Skills & Qualifications
• 7+ years of experience in SRE/DevOps roles.
• 3+ years of team management experience.
• Experience with cloud platforms (AWS/Azure/GCP).
• Knowledge of CI/CD tools (Jenkins, GitLab CI).
• Experience with Docker and Kubernetes.
• Scripting skills (Python, Bash).
• Knowledge of Monitoring tools (Prometheus, Grafana, ELK).
Preferred Qualifications
• Experience with microservices.
• Cloud certifications are a plus.
• Robust problem-solving skills.