17 Sep
|
Xebia IT Architects
|
Gurugram
17 Sep
Xebia IT Architects
Gurugram
L1 Monitoring & AIOps Analyst
Experience: 1-3 years relevant experience
Role Category: Monitoring / AIOps Operations
Primary Objective: 24x7 monitoring and first-line incident response, with a strong focus on alert quality, event correlation, AIOps improvement and continuous reduction of manual operational effort.
Role Purpose
Support 24x7 monitoring across infrastructure, platforms and applications while continuously improving signal quality, alert correlation, triage effectiveness, and operational automation. The role will operate within defined runbooks and ITSM processes while identifying opportunities to reduce noise, improve monitoring and contribute to the AIOps/automation backlog.
Key Responsibilities
- Perform 24x7 monitoring of infrastructure, platforms and applications using Datadog, Splunk, AppDynamics, ScienceLogic or equivalent tools.
- Analyse and triage alerts & events, identify actionable vs. non-actionable events, and perform event correlation/deduplication and noise reduction.
- Create, enrich, prioritize, and route incidents through ServiceNow/PagerDuty as per defined processes.
- Perform initial diagnostics using logs, metrics, dashboards, knowledge articles, and runbooks.
- Identify P1/P2 indicators and trigger the appropriate escalation/Major Incident Management process.
- Monitor alert patterns and recurring incidents to identify opportunities for AIOps, automation and monitoring improvements.
- Support alert-threshold tuning, dashboard validation, monitoring-health checks and observability improvements.
- Contribute operational insights to the AIOps/automation improvement backlog, including opportunities for automated triage, diagnostics, and runbook automation.
- Maintain accurate incident timelines, shift handovers, and operational documentation to support MTTD, MTTA and MTTR measurement.
- Work collaboratively with L2/SRE, platform and automation teams to improve reliability and reduce repetitive manual activities.
Must-Have Skills & Experience
- 1–3 years of experience in monitoring, NOC, production operations, or IT operations.
- Hands-on experience with infrastructure/application monitoring and alert/event triage.
- Exposure to ServiceNow, PagerDuty or equivalent ITSM/on-call tools.
- Ability to interpret basic logs, metrics, alerts, and service dependencies and follow runbooks.
- Understanding of event correlation, alert noise reduction, and incident prioritization.
- Basic understanding of cloud environments (AWS/Azure/GCP).
- Strong incident communication, documentation, and shift-handover discipline.
- Willingness to work in 24x7 rotational shifts.
Good-to-Have / Preferred
- Experience with Datadog, Splunk, AppDynamics, ScienceLogic or Moogsoft.
- Basic Python/scripting knowledge.
- Exposure to AIOps, event correlation, anomaly detection, or automated incident triage.
- Understanding of runbook automation/self-healing concepts.
- ITIL Foundation certification.
- Experience identifying recurring operational issues and converting them into automation/continuous-improvement opportunities.
- Strong signal-vs-noise judgment and ability to remain calm during high-alert volumes.
Communication & Business Writing
- Positive verbal and written communication skills.
- Ability to clearly document incidents, observations, troubleshooting steps, and improvement opportunities.
- Ability to communicate effectively with L2/SRE and engineering teams during escalations.
📌 Monitoring & AIOps Analyst (Gurugram)
🏢 Xebia IT Architects
📍 Gurugram