- 8 to 11 years in Site Reliability Engineering, DevOps, Platform Engineering, or Production Engineering.
- Experience supporting enterprise-scale production environments.
- Proven ownership of services with 99.9%+ uptime commitments.
- Demonstrated experience creating and managing SLOs and Error Budgets.
- Deep troubleshooting expertise in Kubernetes-based production systems.
- Experience handling Sev-1 and Sev-2 incidents.
- Solid understanding of distributed systems and microservices architectures.
- Hands-on cloud platform administration experience.
- Defining SLOs & SLIs
- Have experience in Service mesh
- Exposure to observability tool Datadog (good to have)
Deal Breaker Skill
Reliability Analysis
Mandatory Skills
Site Reliability Engineering| DevOps| Reliability Engineering