16 Sep
|
AxisMaxlife
|
India
About Axis Max Life
Axis Max Life Insurance Limited, formerly known as Max Life Insurance Company Ltd., is a joint venture between Max Financial Services Limited (“MFSL”) and Axis Bank Limited.
Axis Max Life Insurance offers comprehensive protection and long-term savings life insurance solutions through its multi-channel distribution, including agency and third-party distribution partners. It has built its operations over two decades through a need-based sales process, a customer-centric approach to engagement and service delivery, and a well-trained human capital.
Axis Max Life has been consistently ranked among the best workplaces by the GPTW Institute, reflecting its commitment to creating a positive and empowering work environment.
#ComeAsYouAre LGBTQIA+ and PwD candidates of all ages are encouraged to apply
Job Description:
Job Description
Position
Vice President – AIOps and System Reliability Engineering
Incumbent
N/A
Department
Digital Technology
Function
IT Services
Reporting to
CVP and Head of IT Infrastructure & Digital Operations
Band
3
Location
Gurugram
Team Size (D / I)
5-7 / 50-80
JOB SUMMARY:
The Vice President – AIOps and System Reliability Engineering is a visionary technology leader responsible for driving the strategy, automation, and transformation of the enterprise IT operations landscape. This expanded role oversees digital infrastructure, reliability engineering, and cutting‑edge adoption of Agentic AI and autonomous operations to deliver a self‑healing, predictive, and highly automated IT environment.
This leader is accountable for significantly reducing manual support dependencies by implementing AI‑driven automation frameworks that eliminate L1 support (100%) and reduce L2 support effort by at least 50%, while ensuring world‑class performance, reliability, and security across IT services.
Key Responsibilities, Relationships & Measures of Success:
-
Strategic & Autonomous Operations Leadership
o Define and execute a long-term strategy for Agentic AI, autonomous operations, and AI driven service management.
o Build and operationalize an enterprise-wide framework for Autonomous IT Operations (AIOps), ensuring seamless integration with infrastructure, cloud, and SRE functions.
o Lead the implementation of AI agents, decisioning engines, and self-healing automation across service operations.
- Operational Excellence & Automation Transformation
o Achieve 100% elimination of L1 support through predictive automation, intelligent routing, and autonomous resolution workflows.
o Deliver 50% reduction in L2 support workload through AI based diagnostics, automated remediations, and knowledge orchestration.
o Oversee implementation of AI driven monitoring, anomaly detection, auto triaging, and automated incident remediation.
o Optimize IT operations through AIOps, observability platforms, and closed loop automation.
- Infrastructure Engineering & SRE Leadership
o Oversee all DC/DR operations including servers, storage, databases, networking, and hybrid cloud infrastructure.
o Strengthen and scale SRE practices including SLOs, SLIs, error budgets, and automated reliability engineering.
o Build autonomous SRE capabilities through AI driven reliability tooling and prevention first engineering.
- Innovation & Technology Modernization
o Identify, evaluate, and implement emerging technologies including Agentic AI, GenAI copilots, predictive operations, and advanced automation frameworks.
o Lead modernization of CI/CD, IaC, and DevSecOps with embedded AI and smart orchestration.
o Build a center of excellence for autonomous operations and AI first service engineering.
- Incident, Problem & Change Management Automation
o Deploy automated playbooks, AI guided root cause analysis, and recommendation engines.
o Implement self-service and conversational AI capabilities across ITSM platforms.
o Ensure proactive detection (MTTD < 5 minutes) and rapid recovery (MTTR < 1 hour) through automation.
- People & Vendor Leadership
o Lead and mentor engineering, SRE, and automation teams with a robust culture of innovation and accountability.
o Manage strategic vendor partnerships for AI, observability, infrastructure, and automation technologies.
- Stakeholder Engagement
o Communicate operational health, transformation progress, AI impact, and risk posture to senior leadership.
o Partner with cybersecurity, product, and enterprise architecture teams to ensure secure and compliant autonomous operations.
Required Skills & Experience:
-
Bachelor’s degree in engineering, Computer Science, or related field (B.E./BTech preferred).
- 14+ years of progressive experience in infrastructure engineering, SRE, DevOps, and IT operations automation.
- Strong hands-on experience with AIOps platforms, Agentic AI models, and autonomous operations frameworks.
- Proven background in large scale IT modernization, observability, and reliability engineering.
- Expertise in cloud operations (AWS/Azure/OCI), Kubernetes, container orchestration, and IaC.
- Deep understanding of AI/ML, automation platforms, scripting (Python/Shell), and integration pipelines.
- Experience with ITSM platforms, incident automation, and workflow orchestration (ServiceNow preferred).
- Strong leadership capabilities with experience in driving major automation transformations.
- Solid hands-on experience with AIOps platforms, Agentic AI models, and autonomous operations frameworks.
- Proven expertise in managing large-scale, distributed systems with a focus on scalability, reliability, and security.
- Hands-on experience with service management and change management tools (e.g., ServiceNow).
- Deep knowledge of disaster recovery, business continuity, and infrastructure lifecycle management.
- Mastery of DevOps and SRE metrics (DORA, toil reduction) and automation tools (Terraform, CloudFormation, Jenkins, Ansible, etc.).
- Proficiency in security tools (SAST, DAST, container security) and monitoring platforms (Nagios, Dynatrace, SolarWinds).
- Demonstrated ability to lead cross-functional teams, drive change, and deliver results in a fast-paced environment.
Key Performance Indicators (KPIs):
- 100% elimination of L1 support workload through autonomous resolution.
- 50% reduction in L2 dependency via AI driven triage and remediation.
- Deployment downtime reduces by 50%.
- MTTR for SEV1 issues < 1 hour.
- MTTD < 5 minutes through automation.
- IT uptime 99.5%.
- Run cost optimization 10% YoY.
- SLA adherence 95%.
- Automated VAPT across all applications.
- Continuous improvement in system reliability and customer satisfaction scores
- SLA Adherence 95% (critical applications)
- Audit & compliance closure 90% on time
State:
Home Office
Branch:
Gurugram -90C
Department:
Digital Technology
Function:
IT Service Delivery & Operations
Posted On:
03-Feb-2026
📌 Vice President - Site Reliability Engineering (India)
🏢 AxisMaxlife
📍 India