ABOUT THE ROLE
We are hiring an AIOps Engineer to apply machine learning and analytics to IT operations — turning logs, metrics, traces and events into automated anomaly detection, root-cause analysis, incident prediction and self-healing across our client's cloud and platform estate.
WHAT YOU'LL DO
• Build ML models for anomaly detection, event correlation, alert-noise reduction and incident prediction across logs, metrics, traces and events.
• Implement automated root-cause analysis and self-healing / auto-remediation workflows.
• Integrate AIOps with observability stacks (Prometheus, Grafana, OpenTelemetry, ELK, Datadog, Dynatrace, Splunk) and ITSM (ServiceNow).
• Deploy and tune AIOps platforms (Moogsoft, BigPanda, Splunk ITSI) for alert correlation and reduction.
• Build capacity-forecasting and time-series models to pre-empt outages.
• Monitor model accuracy and drift; partner with SRE/on-call teams to reduce MTTR.
SKILLS & REQUIREMENTS
Must-have:
Python, Machine Learning (anomaly detection & time-series),
Observability & Telemetry (Prometheus/Grafana/OpenTelemetry/ELK/Datadog/Dynatrace/Splunk), Event Correlation & Root-Cause Analysis, ITSM (ServiceNow), Cloud Platforms (AWS/Azure/GCP), Incident Management / SRE
Good-to-have:
AIOps platforms (Moogsoft, BigPanda, Splunk ITSI), Kafka / streaming data, log analytics, Kubernetes monitoring, capacity forecasting, scikit-learn / statistical modelling, ChatOps, MLOps
Certifications:
Cloud certification (AWS/Azure/GCP) and/or an observability cert (e.g. Datadog) preferred.
WHAT WE LOOK FOR
• 5+ years in SRE/DevOps/ITOps or ML with hands-on AIOps or ML-for-operations experience.
• Solid Python and applied ML (anomaly detection, forecasting) skills.
• Deep understanding of observability, incident management and reducing MTTR.
How to apply:
Apply via LinkedIn or email your CV to
[email protected]
.
📌 AIOps Engineer (Bengaluru)
🏢 NLightN
📍 Bengaluru