AIOps Support Lead/Manager - Observability (India)

AIOps Support Lead/Manager - Observability (India)

28 Sep
|
Magnet HR Consulting Services
|
India

28 Sep

Magnet HR Consulting Services

India

We are seeking an experienced and dynamic AIOps Support Lead (Manager) to lead a modern, standardized observability and monitoring capability across a diverse enterprise application landscape.

The role will manage a 14-member Tier 1 AIOps Support team, initially supporting approximately 50 applications and scaling to nearly 320 applications over the next two years. The position will own people management, operational processes, triage quality, telemetry/data discipline, and continuous automation initiatives.

The key objective is to transition the organization from reactive, manual incident triage toward proactive and increasingly automated monitoring, supported by a mature OpenTelemetry-based observability framework.

Key Responsibilities :

1. Team &

- Shift Management :
- Lead, hire, coach, develop, and manage a team of 14 Tier 1 AIOps Support Engineers.
- Manage team performance, productivity, capability development, and resource allocation.
- Design and maintain effective 24x7 shift and roster coverage.
- Scale team operations in line with application portfolio growth.
- Establish clear roles, responsibilities, operating standards, and performance expectations.

2. Operations &

- Triage Management :
- Own the quality, consistency, and speed of incident detection, classification, and routing.
- Establish and monitor standards for Mean-Time-to-Triage (MTTT).
- Develop a structured and repeatable framework for onboarding applications into the AIOps monitoring ecosystem.
- Scale application onboarding from approximately 50 to 320 applications.
- Lead major outage and incident response activities.
- Coordinate with application, infrastructure, cloud, and other technical teams during critical incidents.
- Drive Root Cause Analysis (RCA) following major incidents.
- Develop and continuously improve technical runbooks, operating procedures, and knowledge repositories.

3. Observability &

- Data Quality :
- Establish standards for ingesting and managing heterogeneous telemetry and log data.
- Partner with application and technology teams to improve the quality, consistency, and usability of monitoring data.
- Support the transition toward a standardized OpenTelemetry backbone.
- Improve alert classification, correlation, enrichment, and prioritization.
- Reduce alert noise and false positives through continuous monitoring and optimization.
- Establish robust data-quality governance to ensure reliable automated triage.

4. Automation &

- Proactive Monitoring :
- Drive continuous reduction of manual triage activities through automation.
- Identify opportunities for automated alert correlation, event enrichment, incident routing, and early-warning detection.
- Support the transition from reactive monitoring to proactive and predictive operations.
- Prepare the observability environment for future AI-driven and Agentic AI automation.




- Collaborate with engineering and platform teams to identify automation use cases and improve operational efficiency.

5. Incident, Problem &

- Change Management :
- Ensure adherence to ITIL-based Incident, Problem, and Change Management processes.
- Drive effective incident escalation and resolution workflows.
- Ensure appropriate linkage between incidents, problems, changes, configuration items, and business services.
- Maintain alignment with CMDB and service-management practices.
- Monitor adherence to SLA/SLO commitments and operational governance standards.

6. KPI &

- Governance Reporting :
- Define and track operational KPIs and service-quality metrics.
- Prepare regular governance reports for program and technology leadership.
- Monitor metrics including :
- SLA/SLO compliance
- Mean-Time-to-Triage (MTTT)
- Alert-to-incident accuracy
- False-positive rate
- Incident volume and trends
- Automation percentage
- Manual triage reduction
- Application onboarding progress
- Present insights, trends, risks, and improvement opportunities to senior stakeholders.

Technical Requirements :

IT Service Management &

- Governance :
- Strong understanding of ITIL Incident Management lifecycle.
- Knowledge of Problem Management and Change Management.
- Understanding of CMDB concepts and service relationships.
- Strong understanding of SLA/SLO definitions and availability calculations.
- Experience working within structured operational governance frameworks.

Observability &

- Monitoring :
- Hands-on experience with one or more modern monitoring and observability platforms, such as Dynatrace, New Relic, AWS CloudWatch, ManageEngine, Glassbox, or similar enterprise monitoring/observability platforms.
- Additional exposure to OpenTelemetry, distributed tracing, Application Performance Monitoring (APM), log aggregation, metrics and event monitoring, and alert correlation and enrichment.

Infrastructure &

- Networking :
- Good understanding of Virtual Machines (VMs), firewalls, load balancers, containers, Kubernetes, OpenShift / OCP, and enterprise infrastructure and application architecture.

Scripting, APIs &

- Data Formats :
- Working knowledge of Unix/Linux shell scripting and exposure to Windows batch scripting.
- Understanding of JSON and XML data formats.
- Experience with API testing and troubleshooting tools such as Postman, SOAP UI, or similar API testing tools.

Cloud &

- Security :
- Understanding of cloud fundamentals including compute, storage, networking, and cloud monitoring.




- Basic understanding of security technologies and protocols including TLS, SSL, authentication tokens, and secrets management.
- Awareness of data infrastructure concepts such as data lineage, data latency, and data quality.

Leadership &

- Behavioral Competencies :

Team Leadership :

- Proven experience managing and scaling technical support or operations teams.
- Strong people-management, coaching, mentoring, and performance-management skills.
- Ability to manage teams operating in 24x7 support environments.

Scale-Up &

- Transformation :
- Demonstrated experience scaling support operations during periods of significant application or business growth.
- Ability to establish standardized processes while maintaining operational flexibility.

Analytical &

- Problem-Solving Skills :
- Ability to analyze fragmented and high-volume alert data.
- Strong capability to convert noisy monitoring signals into actionable insights.
- Ability to develop effective technical runbooks and operational procedures.

Adaptability :

- Comfortable operating in complex and ambiguous environments.
- Ability to dynamically prioritize applications and support requirements across different lifecycle stages, including Invest, Tolerate, Retire, and Migrate.

Stakeholder Management :

- Strong communication and collaboration skills.
- Ability to work effectively with application owners, infrastructure teams, cloud teams, engineering teams, service-management teams, and senior leadership.

Nice-to-Have Skills :

- Familiarity with AI and GenAI architectures.
- Knowledge of Prompt Engineering, Knowledge Graphs, and Retrieval-Augmented Generation (RAG).
- Experience with cloud-native observability frameworks on AWS, Microsoft Azure, and Google Cloud Platform (GCP).
- Exposure to Agentic AI concepts.
- Experience with AI-driven auto-triage or intelligent automation solutions.
- Familiarity with predictive monitoring and AIOps platforms.

Key Success Measures :

- Success in this role will be demonstrated through :
- A high-performing 14-member Tier 1 AIOps Support team delivering reliable and SLA-compliant triage.
- Successful expansion of monitoring and triage coverage from 50 to approximately 320 applications.
- Consistent improvement in Mean-Time-to-Triage (MTTT) and incident-routing accuracy.
- Reduction in false positives and alert noise.
- Progressive reduction in manual and reactive triage through automation.
- Increased adoption of proactive and early-warning monitoring.
- A standardized and high-quality OpenTelemetry-based telemetry backbone.
- Strong governance of SLA/SLO, alert, incident, and operational-quality metrics.
- Comprehensive runbooks and standardized operational processes.
- A monitoring environment that is ready for future AI-driven and autonomous support capabilities.

📌 AIOps Support Lead/Manager - Observability (India)
🏢 Magnet HR Consulting Services
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: aiops support lead/manager - observability (india) / india