ABOUT
Owns production operations for LLM and agentic workloads — serving, cost, and observability for a fundamentally less predictable class of system than classical ML.
KEY RESPONSIBILITIES
Own production serving and scaling for LLM/agentic workloads (inference infra, load balancing, caching)
Monitor and control inference cost — token usage, retry/loop cost, model routing decisions
Build observability for LLM-specific failure modes: hallucination rate, latency spikes, prompt drift
Manage model/version rollout strategy (canary releases, fallback models, A/B testing)
Own incident response for LLM/agent production issues
Partner with GenAI Engineers and Agentic AI Architects on production-readiness reviews
Explain token-cost dynamics to client finance/business stakeholders
Collaborate closely with GenAI Engineers without needing a hard line between build and run
Support the practice in setting cost governance policy for LLM workloads
REQUIREMENTS & SKILLS
4–6 yrs platform/MLOps engineering with hands-on LLM/GenAI production experience
Deep understanding of LLM inference economics — token costs, batching, caching, model routing
Experience with LLM observability tooling (tracing, eval pipelines, prompt/version management)
Familiarity with multiple model hosting platforms and their cost/performance tradeoffs — Microsoft Azure AI Foundry, AWS Bedrock, and Google Vertex AI, plus self-hosted open-source options (vLLM, TGI) as a positive-to-have
Experience building canary/rollback strategies for probabilistic systems
Comfortable with the higher unpredictability of agentic workloads vs. classical ML serving
Cost-conscious communicator — can explain a token-cost blowup to a client's finance stakeholder
Collaborates closely with GenAI Engineers without needing a hard line between “build” and “run”
Calm under pressure during live incidents affecting client-facing systems
📌 Forward Deployed Engineer (India)
🏢 Systems
📍 India