05 Sep
|
Cloudxtreme
|
Hyderabad
05 Sep
Cloudxtreme
Hyderabad
Role & responsibilities
We are looking for an experienced Site Reliability Engineering Lead (Application SRE Lead) to drive operational excellence, reliability, observability, automation, and service resilience for mission-critical authentication and enterprise applications.
The ideal candidate will have deep expertise in Application SRE, Production Support, Cloud Operations, Incident Management, Authentication Platforms, Observability, Microservices, and DevOps practices while leading teams operating in highly complex, large-scale, globally distributed environments.
This role requires a robust technical leader who spends approximately 80% of time contributing directly to delivery, technical troubleshooting, incident management, automation, observability, deployments, and platform reliability, while dedicating 20% of time to team leadership, stakeholder management, mentoring, governance,
and service improvement initiatives.
The individual will work closely with Engineering, Product, Infrastructure, Security, and Client stakeholders to ensure application availability, reliability, security, and performance meet business objectives.
Mandatory Skills
1.Application SRE, Observability, Incident Management & Reliability Engineering
2.Authentication Platforms (OAuth, OIDC, SAML, MFA, Okta, Transmit, Any)
3.Kubernetes, Docker, Cloud Platforms & Microservices (Any cloud)
4.CI/CD, Release Engineering & Production Deployments (Any)
5.Database & Middleware Support (MongoDB, PostgreSQL, Kafka, Any)
6.Leadership, Delivery Management & Service Excellence
📌 Production Support Lead (Hyderabad)
🏢 Cloudxtreme
📍 Hyderabad