Pune, Maharashtra
Job Summary
We're looking for a Lead / Manager of AI Observability to own how we see, measure, and trust our AI systems once they're live in production. This role sits at the intersection of MLOps, platform engineering, and applied AI, and is responsible for building the monitoring, tracing, evaluation, and alerting infrastructure that gives engineering, product, and leadership real-time visibility into model and agent behavior, quality, cost, and risk. You'll define the observability strategy for LLM-powered and traditional ML systems in production, lead a small team (or embedded function) of engineers, and partner closely with ML, platform, and reliability teams to catch issues before customers do.
Key Responsibilities
Own the end-to-end observability strategy for AI systems in production, including tracing, logging, metrics, evaluation pipelines, and alerting.
Design and build systems to monitor model/agent quality in production: accuracy drift, hallucination rate, latency, cost per request, token usage, and task success rate.
Establish golden signals and SLOs for AI systems, distinct from traditional infra SLOs (e.g., output quality, safety, groundedness, factuality).
Build or integrate tracing across multi-step/agentic workflows so failures can be root-caused across prompts, tool calls, retrieval steps, and model versions.
Stand up automated evaluation frameworks (offline and online/production evals) to continuously score live traffic and catch regressions after model, prompt, or data updates.
Partner with ML/platform engineering to instrument new models and features with observability hooks before they reach production.
Define and drive incident response processes specific to AI failures (silent quality degradation, drift, prompt injection, unsafe outputs) — not just uptime.
Build dashboards and reporting for engineering, product, and executive stakeholders on production AI health, cost, and risk posture.
Lead, mentor,
and grow a team of engineers focused on observability tooling, or act as the technical lead embedded across ML/platform teams.
Evaluate, select, and manage the observability toolchain (build vs. buy) across tracing, evals, monitoring, and cost-tracking platforms.
Partner with security, compliance, and legal on auditability, data retention, and responsible-AI monitoring requirements.
Skill Requirements
+ years in software/ML engineering, with 3+ years focused on observability, monitoring, reliability, or MLOps.
Hands-on experience running AI/ML systems in production, including at least one LLM-based or generative AI system at scale.
Solid understanding of distributed tracing, structured logging, and metrics pipelines (e.g., OpenTelemetry-style concepts), applied to AI/agentic workflows.
Experience designing evaluation frameworks for generative AI (offline benchmarks, online/production evals, human-in-the-loop review).
Solid grasp of the unique failure modes of production AI: drift, hallucination, prompt injection, latency/cost blowups, silent quality regressions.
Track record of building or leading a team, or serving as a technical lead across cross-functional engineering groups.
Proficiency in at least one major programming language (Python, Go, or similar) and comfort working across the ML/platform stack.
Excellent cross-functional communication — able to translate observability data into decisions for engineers, product managers, and executives.
Other Requirements
Experience with vector databases, RAG pipelines, or agentic frameworks in production.
Background in SRE/DevOps prior to moving into ML/AI observability.
Familiarity with responsible AI, model risk management, or AI governance frameworks.
Experience presenting observability/risk posture to executive or board-level audiences.
#body.unify div.unify-button-container .unify-apply-now: focus, #body.unify div.unify-button-container .unify-apply-#body.unify div.unify-button-container .unify-apply-now: focus, #body.unify div.unify-button-container .unify-apply-
📌 Senior Technical Lead (India)
🏢 HCLTech
📍 India