31 Jul
|
LA Consultancy
|
India
31 Jul
LA Consultancy
India
Qualifications :
BE, BTech, MCA, or equivalent degree in Computer Science, IT, or a related field with 8+ years of experience in DevOps, SRE, Cloud Operations, Platform Engineering, or Production Engineering, with demonstrated technical leadership responsibilities.
Technical Requirements :
- Strong experience supporting highly available SaaS or cloud platforms in a 247 production environment.
- Hands-on experience with one or more major cloud providers : AWS, Google Cloud Platform, or Microsoft Azure.
- Strong experience with Kubernetes and containerized production environments.
- Experience with Infrastructure as Code, preferably Terraform, along with Jenkins, GitOps, and automation practices.
- Solid Linux, networking, troubleshooting, and distributed systems fundamentals.
- Deep experience with observability, Incident Management, On-Call operations, RCA, SLA, SLO, and production reliability practices.
Leadership and Operational Skills :
- Proven ability to define, manage, and improve operational KPIs and translate trends into measurable actions and outcomes.
- Experience leading complex customer-critical incidents, coordinating multiple technical teams, and driving issues through resolution.
- Strong ability to influence cross-functional teams and drive accountability without relying on direct authority.
- Proven ability to mentor engineers, improve operational practices, and raise technical and execution standards.
- Excellent communication skills with the ability to engage technical teams, cross-functional stakeholders, and leadership during critical situations.
- Strong ownership mindset with a focus on customer outcomes, operational excellence, and continuous improvement.
Core Responsibilities :
- Lead 247 follow-the-sun operations, serving as the front line for infrastructure alerts and Incident Management, with effective On-Call coverage, rapid response, service restoration, escalation, and cross-region handoffs.
- Own and drive Customer Experience and Operational KPIs, including Customer Issue and DOHD ticket aging, response time, backlog burn-down, escalations, routing quality, go-live incidents, and RCA SLA compliance.
- Lead daily operational triage of new, aging, blocked, and escalated work, ensuring clear priorities, ownership, accountability, and timely closure.
📌 Lead DevOps Engineer (India)
🏢 LA Consultancy
📍 India