Role & responsibilities
8 to 11 years in Site Reliability Engineering, DevOps, Platform Engineering, or Production Engineering.
Experience supporting enterprise-scale production environments.
Proven ownership of services with 99.9%+ uptime commitments.
Demonstrated experience creating and managing SLOs and Error Budgets.
Deep troubleshooting expertise in Kubernetes-based production systems.
Experience handling Sev-1 and Sev-2 incidents.
Solid understanding of distributed systems and microservices architectures.
Hands-on cloud platform administration experience.
Defining SLOs & SLIs
Have experience in Service mesh
Exposure to observability tool Datadog (positive to have)
Deal Breaker Skill
Reliability Analysis
Mandatory Skills
Site Reliability Engineering| DevOps| Reliability Engineering