13 Aug
|
Elabs Infotech
|
Bengaluru
13 Aug
Elabs Infotech
Bengaluru
Experience: 13 - 16 Years
Location: Bellandur, Bangalore
Shift: India / US Shift
Role: Principal Architect
About the Role
We are looking for an experienced AI Observability Principal Architect to define and drive enterprise-wide observability, SRE, AIOps, and intelligent automation strategies. The ideal candidate will have solid expertise in Grafana, OpenTelemetry, SRE practices, AIOps, event management, automation, and infrastructure observability.
The role requires an architect who can establish observability standards, implement modern telemetry frameworks, drive AI-driven monitoring initiatives, and enable automated incident management and self-healing across enterprise platforms.
Key Responsibilities
- Define and implement the enterprise observability strategy, architecture, standards, and best practices.
- Design and implement observability solutions covering logs, metrics, and distributed traces using OpenTelemetry.
- Develop enterprise dashboards, alerting frameworks, and operational views using Grafana.
- Establish SLIs, SLOs, Error Budgets, and Critical User Journeys (CUJs) in collaboration with SRE and application teams.
- Lead AIOps and Event Management initiatives including event correlation, anomaly detection, predictive analytics, and alert-noise reduction.
- Design and implement incident automation, intelligent routing, remediation, and self-healing capabilities.
- Integrate observability platforms with ITSM platforms, preferably ServiceNow ITOM.
- Drive automation using PowerShell and Ansible.
- Collaborate with SRE, Infrastructure, Application, Cloud, and Service Management teams.
- Build executive-level and operational dashboards for service health, performance, availability, and reliability.
- Provide technical leadership during major incidents and drive continuous service improvement.
- Evaluate and implement AI/ML-driven observability and monitoring solutions.
- Support automation initiatives using Power Apps and Power Automate where applicable.
Primary Skills
- Grafana Observability, Dashboarding & Alerting
- OpenTelemetry Logs, Metrics & Distributed Tracing
- SRE Practices – SLI, SLO, Error Budgets, CUJs
- AIOps & Event Management
- Splunk / SolarWinds
- Azure Observability / Azure Monitoring
- Incident Management & Automation
- PowerShell
- Ansible
- AI/ML-driven Monitoring and Predictive Analytics
- Strong understanding of Windows/Linux infrastructure, networks, and firewalls
Secondary / Nice-to-Have Skills
- ServiceNow ITOM
- Service Mapping
- IntegrationHub / integrations
- Identification and Reconciliation Engine (IRE)
- AIOps
- ServiceNow FSM
- ServiceNow ITAM / HAM
- Power Apps
- Power Automate
- Cloud and platform observability
- Predictive analytics
- AI-driven monitoring and anomaly detection
Required Experience
- 13+ years of experience across Observability, SRE, IT Operations, Platform Engineering, Monitoring, or AIOps.
- Proven experience designing and implementing enterprise observability architectures.
- Strong hands-on and architectural expertise in Grafana and OpenTelemetry.
- Strong understanding of monitoring, telemetry, alerting, incident management, and automation.
- Experience implementing SRE principles, including SLIs, SLOs, error budgets, and reliability engineering.
- Experience integrating observability platforms with ITSM tools, preferably ServiceNow.
- Strong leadership and stakeholder management skills with the ability to work across infrastructure, application, SRE, and service management teams.
Preferred Profile
Candidates with experience in the following areas will be preferred:
- Enterprise-scale AIOps transformation
- AI/ML-based anomaly detection and predictive monitoring
- ServiceNow ITOM / AIOps
- Automated remediation and self-healing infrastructure
- Azure observability
- Large-scale Grafana and OpenTelemetry implementations
- Major Incident Management and Continuous Service Improvement
📌 AI Observability & Monitoring Engineer (Bengaluru)
🏢 Elabs Infotech
📍 Bengaluru