AI Ops Engineer (Chennai)

AI Ops Engineer (Chennai)

08 Aug
|
Mastech Digital
|
Chennai

08 Aug

Mastech Digital

Chennai

AIOps Engineer

DevOps Artificial Intelligence Site Reliability Engineering

Function

Platform Engineering & AI Operations

Experience

4 – 8 Years

Python

Kubernetes / EKS

Prometheus + Grafana

ML Ops

Terraform / IaC

CI/CD

About the Role

We are looking for a seasoned AIOps Engineer who sits at the intersection of platform reliability, AI/ML operations, and intelligent automation. You will design and operate AI-driven observability pipelines, build self-healing infrastructure, and embed machine-learning models into DevOps toolchains — turning telemetry noise into actionable intelligence at scale.

As a trusted member of the Platform Engineering team, you will own the full lifecycle of AIOps tooling: from anomaly-detection models and intelligent alerting to predictive capacity management and LLM-assisted incident response — all running on cloud-native, containerised infrastructure.

Key Responsibilities

AIOps & Intelligent Automation

- Design and operate AI/ML-powered observability platforms (anomaly detection, root-cause analysis, predictive alerting) using tools such as Dynatrace, New Relic, Moogsoft, or open-source equivalents.
- Build and maintain Python-based ML pipelines (scikit-learn, PyTorch, or TensorFlow) for log analytics, metric forecasting, and event correlation.
- Develop intelligent auto-remediation runbooks triggered by model-detected incidents, reducing MTTR by 40%+.
- Integrate LLM-assisted copilots (OpenAI / Azure OpenAI) into incident management workflows for automated triage and war-room summaries.

DevOps & Platform Engineering

- Own CI/CD pipeline design and maintenance using GitHub Actions, GitLab CI, or Jenkins; enforce shift-left security with SAST/DAST, SBOM, and Trivy scans.
- Provision, scale, and manage Kubernetes clusters (EKS / AKS / GKE) using Helm, Kustomize, and ArgoCD with GitOps workflows.
- Manage Infrastructure-as-Code using Terraform and Ansible across multi-cloud environments (AWS primary, Azure secondary).
- Champion containerisation best practices: image hardening, multi-stage builds, and supply-chain security with Cosign/Sigstore.

Site Reliability Engineering

- Define, implement, and operationalise SLOs, SLIs, and error-budget burn policies aligned to business objectives.
- Architect and maintain a four-layer observability stack: metrics (Prometheus/Thanos), logs (OpenSearch/Loki), traces (Jaeger/Tempo),



and dashboards (Grafana).
- Lead blameless post-mortem processes and own the incident management lifecycle end-to-end.
- Drive capacity planning using ML-based demand forecasting models integrated with KEDA autoscaling policies.

Python Engineering & Tooling

- Write production-grade Python automation: alerting integrations, CMDB reconciliation, drift-detection daemons, and Slack/PagerDuty bots.
- Build internal CLI tools and REST APIs (FastAPI / Flask) to expose AIOps insights to engineering teams.
- Maintain data pipelines feeding telemetry into feature stores and model retraining loops.

Required Qualifications

Experience & Education

- 4–8 years of hands-on experience in DevOps, SRE, or Platform Engineering roles.
- 2+ years in an AIOps or MLOps capacity — deploying and monitoring ML models in production.
- Bachelor's or Master's degree in Computer Science, Information Technology, or equivalent.

Technical Skills — Must Have

- Python (3.x): proficient in scripting, automation, REST APIs, ML libraries, and unit testing (pytest).
- Kubernetes: workload design, RBAC, HPA/KEDA, network policies, and multi-tenancy.
- Observability stack: Prometheus, Grafana, Alertmanager, OpenTelemetry, and distributed tracing.
- CI/CD: GitHub Actions or GitLab CI — pipeline authoring, environment promotion, and gate policies.
- IaC: Terraform (modules, state management, remote backends) and Helm chart authoring.
- Cloud: Strong hands-on Solution Architect experience on AWS and Azure — production-grade infrastructure design, multi-cloud networking, security, and cost governance.
- Incident Management: PagerDuty, Opsgenie, or VictorOps integrated with AIOps runbooks.

Technical Skills — Robust Advantage

- ML frameworks: scikit-learn, Prophet, or PyTorch for time-series anomaly detection.
- Kafka / NATS JetStream for high-throughput telemetry streaming.
- Service mesh: Istio or Linkerd for traffic observability and canary rollouts.
- Security tooling: Trivy, Falco,



OPA/Gatekeeper, OWASP LLM Top 10 awareness.

LLM / SLM Optimization

- Hands-on experience deploying and optimising Large Language Models (GPT-4o, Claude, Llama 3) and Small Language Models (Phi-3, Mistral, Gemma) in production AIOps pipelines.
- Model quantisation techniques: GGUF/GGML, GPTQ, AWQ — reducing inference footprint by 2–4 without significant accuracy loss.
- Prompt engineering and chain-of-thought optimisation for ops-domain tasks such as log summarisation, RCA generation, and runbook synthesis.
- Retrieval-Augmented Generation (RAG): building domain-specific vector stores (pgvector, MongoDB Atlas Vector Search, Pinecone) backed by incident knowledge bases.
- LLM inference serving: vLLM, Ollama, or TGI (Text Generation Inference) on Kubernetes — autoscaling with KEDA based on token throughput.
- Fine-tuning SLMs on internal ops datasets using LoRA / QLoRA for domain-specific intent classification (alert triage, ticket routing, anomaly labelling).
- LLM observability: token usage tracking, latency p99 SLOs, hallucination guardrails, and prompt injection detection using tools such as Langfuse or LangSmith.
- Context-window management and cost optimisation strategies: prompt compression, semantic caching, and batching via Anthropic / OpenAI batch APIs.
- Model routing and A/B evaluation frameworks — switching between frontier LLMs and local SLMs based on latency, cost, and task complexity signals.

Behavioural Competencies

- Ownership mindset — you treat production systems as if they were your own product.
- Data-driven decision-making — you instrument first, optimise second.
- Collaborative problem-solver — comfortable pairing with data science, security, and product teams.
- Strong written communication — able to author RCAs, design docs, and architecture decision records.
- Growth orientation — curious about emerging AI tooling and proactive in knowledge sharing.

Nice to Have

- Experience on high-scale consumer platforms (50M+ users, sub-200ms p99 SLOs).
- Contributions to open-source AIOps or observability projects.
- Familiarity with MCP (Model Context Protocol) agentic architectures.
- AWS / GCP / Azure certifications at Professional or Specialty level.
- Knowledge of FinOps tooling (Kubecost, AWS Cost Explorer API integration).

-

📌 AI Ops Engineer (Chennai)
🏢 Mastech Digital
📍 Chennai

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: ai ops engineer (chennai) / chennai