- Manage P1/P2 incidents, escalation, and resolution
- Drive RCA, problem management, and preventive fixes
- Improve MTTR, MTTD, system reliability
2. Team Leadership
- Lead and manage SRE / Ops team.
- Allocate work, run scrum/standups, track deliverables
- Mentor engineers and ensure operational discipline
3. Cloud & Platform
- Strong experience in AWS-based environments
- Understanding of Kubernetes (EKS) and microservices architecture
- Work with CI/CD pipelines (Jenkins mandatory)
4. Infrastructure & Automation
- Experience with Terraform / CloudFormation
- Exposure to Ansible / scripting (Python/Shell)
- Understanding of infra provisioning and debugging issues