AIOps Support Lead, Manager (India)

AIOps Support Lead, Manager (India)

29 Aug
|
Talent Socio
|
India

29 Aug

Talent Socio

India

About the Role

Our client's AIOps program is building a modern, standardized observability and monitoring capability across a diverse application landscape. We are starting with Tier 1 (L1) triage support for 50 applications and scaling to roughly 320 applications within two years. The AIOps Support Lead will manage a team of 14 AIOps Support Engineers, own the quality of day-to-day manual triage, and drive the transition toward automated, proactive triage as the Open Telemetry backbone matures. This role owns the people, process, and data-quality discipline that makes reliable triage possible it does not own the Agentic automation layer itself but prepares the team and the data for it.

What You'll Do

Manage the Tier 1 team: directly manage a team of 14 AIOps Support Engineers performing manual triage of alarms and alerts across a diverse, 50-to-320-application portfolio including hiring, coaching, scheduling, and performance management.

Own triage quality and speed: set and monitor standards for how quickly and accurately the team detects, classifies, and routes incidents, and drive continuous improvement in mean-time-to-triage.

Drive data stewardship: partner with application teams to standardize alarm and alert data across heterogeneous log aggregation tools (Dynatrace, New Relic, Manage Engine, Glass box, and others) into a clean, consistent telemetry backbone built on Open Telemetry.

Manage the reactive-to-proactive shift: reduce reliance on reactive, manual triage over time by improving alert quality, correlation, and early-warning signals laying the groundwork for future automated and Agentic triage.

Navigate a diverse, moving application landscape: support applications spanning different technology stacks and different architecture dispositions (Invest, Tolerate, Retire, Migrate), reprioritizing team focus as the portfolio shifts.

Coordinate onboarding of new apps:



run a repeatable process for bringing new applications into Tier 1 coverage as the program scales from 50 to 320 applications, including support group and application owner mapping.

Manage stakeholders: act as the primary point of contact for support groups, application owners, and AIOps program leadership on Tier 1 status, incidents, and data-quality issues.

Manage shift/roster coverage: ensure the team of 14 provides consistent triage coverage across required hours as the application count grows.

Report on outcomes: track and report team KPIs triage time, alert-to-incident accuracy, false-positive rates, coverage growth to program leadership.

What We're Looking For

Experience: 7+ years in application/production support (L1/L1.5/L2) or site reliability, with 2+ years directly managing a technical support team.

Hybrid environment expertise: proven experience supporting applications across both on-premises and cloud environments, with exposure to modern microservices architectures.

Observability tooling: hands-on experience with monitoring and observability platforms such as Dynatrace, New Relic, AWS Cloud Watch, Manage Engine, or Glassbox; working knowledge of Open Telemetry and distributed tracing concepts.

ITIL discipline: robust grounding in incident, problem, and change management practices, with Service Now or Jira ticket management experience.

Technical range: comfortable with Linux and Windows troubleshooting, basic networking (TCP/IP, DNS, HTTP/HTTPS, SSL, load balancers), SQL/database query analysis, and API/integration troubleshooting.

People management: demonstrated ability to hire, coach,



and retain a team of 10+ technical support staff through a period of significant scale-up (5x application coverage growth).

Analytical mindset: able to turn noisy, inconsistent alert data into clear, actionable insight, and to build repeatable frameworks rather than one-off fixes.

Comfort with ambiguity: willing to support a moving target a portfolio spanning Invest, Tolerate, Retire, and Migrate applications and to adapt priorities as the program evolves.

Must-Have Skills

Incident Management Lifecycle, and working knowledge of Problem, Change Request, and Service Request concepts (ITIL)

CMDB concepts and their use in incident and asset traceability

Hands-on experience with log aggregation technologies (Dynatrace, New Relic, Manage Engine, Glass box, or similar)

Working knowledge of JSON and XML, and basic file/task automation

Understanding of IT infrastructure and basic networking: VMs, firewalls, load balancers, containers, Open Shift (OCP), Kubernetes

Unix Shell scripting; Windows batch file creation

Basic cloud concepts: compute, storage, and security fundamentals

Security fundamentals: TLS, SSL, tokens, and secret management

Familiarity with API gateways and API testing toolkits (Postman, SOAP UI, or similar)

Outage response management and experience leading cross-functional coordination during major incidents

Ability to drive Root Cause Analyses (RCAs) and build reusable knowledge articles/runbooks

SLA/SLO management and reporting, including availability calculation

Working knowledge of data concepts: data latency, data fragmentation, data lineage, and data marts Nice-to-Have Skills

Familiarity with AI concepts such as prompt engineering, knowledge graphs, and Retrieval-Augmented Generation (RAG)

Experience with cloud-native observability on AWS, Azure, or GCP

Exposure to Agentic AI or automation-driven triage tooling

📌 AIOps Support Lead, Manager (India)
🏢 Talent Socio
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: aiops support lead, manager (india) / india