Your key responsibilities
- Implement AIOps capabilities including anomaly detection, event correlation, alert noise reduction, incident prioritization, and automated remediation.
- Design and maintain observability solutions using logs, metrics, traces, and events.
- Build telemetry pipelines using OpenTelemetry and related technologies.
- Develop automation and self-healing workflows using Python, PowerShell, Bash, or Go.
- Implement AI-powered operational assistants, incident triage, and RAG-based knowledge solutions.
- Integrate observability, CI/CD, ITSM, and cloud platforms to improve operational intelligence.
- Support SRE practices including SLIs, SLOs, incident management, and reliability improvements.
Skills and attributes for success
Required Skills
- AIOps and Observability Platforms (Dynatrace, Datadog, Splunk,
Current Relic, Azure Monitor, Prometheus/Grafana)
- Microsoft Azure and/or AWS
- Kubernetes and Containers
- OpenTelemetry and telemetry pipelines
- Python and automation scripting
- GitHub Actions and/or Azure DevOps
- Incident Management and SRE fundamentals
Good To Have
- Semantic Kernel, LangGraph, LangChain
- Azure AI Foundry or Azure OpenAI
- Kafka, Elasticsearch/OpenSearch
- Terraform, ArgoCD
- ServiceNow integrations
Education:
Degree: Bachelors degree in Computer Science, Engineering, Information Technology, or a related field, or equivalent practical experience
📌 Cloud AIOps (Chennai)
🏢 EY
📍 Chennai