08 Aug
|
Mastech Digital
|
Chennai
08 Aug
Mastech Digital
Chennai
AIOps Engineer
DevOps Artificial Intelligence Site Reliability Engineering
Function
Platform Engineering & AI Operations
Experience
4 – 8 Years
Python
Kubernetes / EKS
Prometheus + Grafana
ML Ops
Terraform / IaC
CI/CD
About the Role
We are looking for a seasoned AIOps Engineer who sits at the intersection of platform reliability, AI/ML operations, and intelligent automation. You will design and operate AI-driven observability pipelines, build self-healing infrastructure, and embed machine-learning models into DevOps toolchains — turning telemetry noise into actionable intelligence at scale.
As a trusted member of the Platform Engineering team, you will own the full lifecycle of AIOps tooling: from anomaly-detection models and intelligent alerting to predictive capacity management and LLM-assisted incident response — all running on cloud-native, containerised infrastructure.
Key Responsibilities
AIOps & Intelligent Automation
- Design and operate AI/ML-powered observability platforms (anomaly detection, root-cause analysis, predictive alerting) using tools such as Dynatrace, New Relic, Moogsoft, or open-source equivalents.
- Build and maintain Python-based ML pipelines (scikit-learn, PyTorch, or TensorFlow) for log analytics, metric forecasting, and event correlation.
- Develop intelligent auto-remediation runbooks triggered by model-detected incidents, reducing MTTR by 40%+.
- Integrate LLM-assisted copilots (OpenAI / Azure OpenAI) into incident management workflows for automated triage and war-room summaries.
DevOps & Platform Engineering
- Own CI/CD pipeline design and maintenance using GitHub Actions, GitLab CI, or Jenkins; enforce shift-left security with SAST/DAST, SBOM, and Trivy scans.
- Provision, scale, and manage Kubernetes clusters (EKS / AKS / GKE) using Helm, Kustomize, and ArgoCD with GitOps workflows.
- Manage Infrastructure-as-Code using Terraform and Ansible across multi-cloud environments (AWS primary, Azure secondary).
- Champion containerisation best practices: image hardening, multi-stage builds, and supply-chain security with Cosign/Sigstore.
Site Reliability Engineering
- Define, implement, and operationalise SLOs, SLIs, and error-budget burn policies aligned to business objectives.
- Architect and maintain a four-layer observability stack: metrics (Prometheus/Thanos), logs (OpenSearch/Loki), traces (Jaeger/Tempo),
and dashboards (Grafana).
- Lead blameless post-mortem processes and own the incident management lifecycle end-to-end.
- Drive capacity planning using ML-based demand forecasting models integrated with KEDA autoscaling policies.
Python Engineering & Tooling
- Write production-grade Python automation: alerting integrations, CMDB reconciliation, drift-detection daemons, and Slack/PagerDuty bots.
- Build internal CLI tools and REST APIs (FastAPI / Flask) to expose AIOps insights to engineering teams.
- Maintain data pipelines feeding telemetry into feature stores and model retraining loops.
Required Qualifications
Experience & Education
- 4–8 years of hands-on experience in DevOps, SRE, or Platform Engineering roles.
- 2+ years in an AIOps or MLOps capacity — deploying and monitoring ML models in production.
- Bachelor's or Master's degree in Computer Science, Information Technology, or equivalent.
Technical Skills — Must Have
- Python (3.x): proficient in scripting, automation, REST APIs, ML libraries, and unit testing (pytest).
- Kubernetes: workload design, RBAC, HPA/KEDA, network policies, and multi-tenancy.
- Observability stack: Prometheus, Grafana, Alertmanager, OpenTelemetry, and distributed tracing.
- CI/CD: GitHub Actions or GitLab CI — pipeline authoring, environment promotion, and gate policies.
- IaC: Terraform (modules, state management, remote backends) and Helm chart authoring.
- Cloud: Strong hands-on Solution Architect experience on AWS and Azure — production-grade infrastructure design, multi-cloud networking, security, and cost governance.
- Incident Management: PagerDuty, Opsgenie, or VictorOps integrated with AIOps runbooks.
Technical Skills — Robust Advantage
- ML frameworks: scikit-learn, Prophet, or PyTorch for time-series anomaly detection.
- Kafka / NATS JetStream for high-throughput telemetry streaming.
- Service mesh: Istio or Linkerd for traffic observability and canary rollouts.
- Security tooling: Trivy, Falco,
OPA/Gatekeeper, OWASP LLM Top 10 awareness.
LLM / SLM Optimization
- Hands-on experience deploying and optimising Large Language Models (GPT-4o, Claude, Llama 3) and Small Language Models (Phi-3, Mistral, Gemma) in production AIOps pipelines.
- Model quantisation techniques: GGUF/GGML, GPTQ, AWQ — reducing inference footprint by 2–4 without significant accuracy loss.
- Prompt engineering and chain-of-thought optimisation for ops-domain tasks such as log summarisation, RCA generation, and runbook synthesis.
- Retrieval-Augmented Generation (RAG): building domain-specific vector stores (pgvector, MongoDB Atlas Vector Search, Pinecone) backed by incident knowledge bases.
- LLM inference serving: vLLM, Ollama, or TGI (Text Generation Inference) on Kubernetes — autoscaling with KEDA based on token throughput.
- Fine-tuning SLMs on internal ops datasets using LoRA / QLoRA for domain-specific intent classification (alert triage, ticket routing, anomaly labelling).
- LLM observability: token usage tracking, latency p99 SLOs, hallucination guardrails, and prompt injection detection using tools such as Langfuse or LangSmith.
- Context-window management and cost optimisation strategies: prompt compression, semantic caching, and batching via Anthropic / OpenAI batch APIs.
- Model routing and A/B evaluation frameworks — switching between frontier LLMs and local SLMs based on latency, cost, and task complexity signals.
Behavioural Competencies
- Ownership mindset — you treat production systems as if they were your own product.
- Data-driven decision-making — you instrument first, optimise second.
- Collaborative problem-solver — comfortable pairing with data science, security, and product teams.
- Strong written communication — able to author RCAs, design docs, and architecture decision records.
- Growth orientation — curious about emerging AI tooling and proactive in knowledge sharing.
Nice to Have
- Experience on high-scale consumer platforms (50M+ users, sub-200ms p99 SLOs).
- Contributions to open-source AIOps or observability projects.
- Familiarity with MCP (Model Context Protocol) agentic architectures.
- AWS / GCP / Azure certifications at Professional or Specialty level.
- Knowledge of FinOps tooling (Kubecost, AWS Cost Explorer API integration).
-
📌 AI Ops Engineer (Chennai)
🏢 Mastech Digital
📍 Chennai