04 Oct
|
The Glove
|
Hyderabad
04 Oct
The Glove
Hyderabad
Observability Architect AIOps & Data Science
Location: [Location]
Experience: 3+ Years
Employment Type: Full-Time
Work Mode: [On-site/Hybrid/Remote]
Position Overview
We are looking for an experienced Observability Architect to join our Platform Engineering and Reliability team, focused on building intelligent systems for AIOps, Data Science, and Operational Intelligence.
The role combines observability, machine learning, AI, telemetry analytics, and platform architecture to develop intelligent solutions for anomaly detection, incident analysis, log correlation, root cause analysis, and automated operational insights.
You will work closely with SRE, Platform Engineering, Data Science, Product, and Software Engineering teams to design scalable observability solutions and shape the architecture and technology strategy for an AI-powered operational intelligence platform.
Key Responsibilities
- Design and develop ML models for anomaly detection, error detection, and root cause analysis using logs, metrics, and telemetry data.
- Develop log correlation and event analysis algorithms across distributed systems.
- Build and maintain knowledge graphs representing system dependencies, service relationships, and failure patterns.
- Develop automated incident analysis pipelines that ingest telemetry, correlate events, and recommend potential root causes.
- Design scalable and resilient telemetry ingestion, aggregation, and analytics pipelines for real-time operational intelligence.
- Define platform standards, technology selections, architecture patterns, and engineering best practices.
- Collaborate with SRE, engineering, and product teams to solve complex enterprise observability challenges using AI.
- Contribute to architecture and design reviews focused on reliability, scalability, security, performance, and automation.
- Mentor junior data scientists and cross-functional teams in ML, AIOps, and observability practices.
- Document ML methodologies, assumptions, performance metrics, and model outcomes.
- Build dashboards and reporting mechanisms for model performance and operational visibility.
- Stay current with emerging technologies in observability, AI workflows, AIOps, and telemetry-driven operational intelligence.
- Optionally develop backend services/APIs using Python, Java, or .NET to integrate ML capabilities with platform services.
Required Skills & Qualifications
- Bachelor's degree in Computer Science, Data Science, Statistics, or a related field, or equivalent experience.
- 3+ years of professional experience in Data Science, Machine Learning Engineering, or Observability Analytics.
- Strong hands-on programming experience with Python.
- Experience developing ML models using TensorFlow, PyTorch, or scikit-learn.
- Strong understanding of NLP techniques for log analysis and time-series anomaly detection.
- Experience with distributed technologies such as Apache Spark, Flink, or Kafka.
- Knowledge of Knowledge Graphs and Graph Neural Networks (GNNs).
- Experience with technologies such as Neo4j, PyG, or DGL is highly desirable.
- Hands-on experience with observability platforms such as:
- ELK / Elastic Stack
- Datadog
- Splunk
- Prometheus
- or equivalent observability platforms.
- Solid understanding of metrics, logs, traces,
dashboards, alerting, SLOs, SLIs, incident management, and operational analytics.
- Exposure to AI-powered operational intelligence, including:
- Agentic AI workflows
- LLM-based operational assistants
- Graph-based dependency mapping
- Automated incident detection and triage
- Root Cause Analysis (RCA)
- Remediation recommendations
- Strong analytical, problem-solving, and communication skills.
Preferred Qualifications
- Master's degree in Machine Learning, Computer Science, Data Science, or a related discipline.
- Experience in AIOps, SRE, IT Operations, or Observability Engineering.
- Experience with causal inference techniques for root cause analysis.
- Published research or contributions to ML, AI, AIOps, or observability open-source projects.
- Experience with Docker, Kubernetes, and CI/CD pipelines.
- Familiarity with incident management platforms and on-call tooling.
- Experience working with microservices and cloud-native architectures.
- Experience with Azure is preferred.
- Understanding of SRE practices, self-healing automation, capacity prediction, and operational decision intelligence.
What You’ll Work On Observability Telemetry ML/AI Correlation Knowledge Graphs RCA Incident Intelligence Automated Remediation
If you are passionate about combining Observability, Machine Learning, GenAI, and AIOps to build intelligent and resilient enterprise platforms, we would like to hear from you.
Keywords
Observability | AIOps | Python | Machine Learning | NLP | Anomaly Detection | Root Cause Analysis | Log Analytics | Time-Series | Knowledge Graph | GNN | Neo4j | PyTorch | TensorFlow | Kafka | Spark | Flink | ELK | Splunk | Datadog | Prometheus | LLM | Agentic AI | SRE | Kubernetes | Azure
📌 Observability Architect AIOps Data Science (Hyderabad)
🏢 The Glove
📍 Hyderabad