04 Oct
|
The Standard India
|
Bengaluru
04 Oct
The Standard India
Bengaluru
AI/ML Platform Engineer
Job Summary The Standard is seeking an AI/ML Platform Engineer IV to stand up the multi-provider model gateway, the model registry and the LLMOps/MLOps lifecycle for platform-owned models on our Data & AI Platform.
In this role, you will define how models are accessed, governed, evaluated and cost-managed across the enterprise routing across Anthropic, OpenAI and open-source models, operating the model registry with risk status, and building evaluation and observability into every model change. You will also set up and operate our LLMOps environment including AI gateway, GPU inference infrastructure: deploying NVIDIA NIM microservices and open-source serving frameworks on NVIDIA Grace Blackwell (GB300) environments to host open-source models on-premises. Your initial assignment is the MCP Gateway task force.
Actuarial pricing and reserving models remain governed by Actuarial's own standards; this role covers platform-owned models only.
This role sits on The Standard's Data & AI Platform team, which owns the enterprise's shared data and AI foundation governed once and consumed by many built on Azure Databricks, Neo4j, Confluent Kafka, Kong, and Anthropic and OpenAI models behind a governed gateway
Ability to work on-site in Bengaluru, India is a requirement of the role.
Mode of Working: This role follows a hybrid work model, with a primary base in Bengaluru. While on-site presence is essential for key engagements, flexibility is offered for remote work based on business needs and team alignment.
Key Responsibilities:
Design and operate the multi-provider model gateway routing, failover, rate limiting, caching and usage attribution (LiteLLM-style patterns) across Anthropic, OpenAI and self-hosted open-source models.
Stand up and operate the MCP/tool registry, AI Gateway a governed catalog of MCP servers and tools with scoped permissions and A2A coordination surfaces.
Set up and operate GPU inference infrastructure on NVIDIA Grace Blackwell (GB300) environments: deploy NVIDIA NIM microservices and open-source serving frameworks such as vLLM, SGLang and NVIDIA Dynamo for distributed serving.
¢ Manage GPU environments end to end: Kubernetes GPU scheduling (GPU Operator, device plugins, MIG/time-slicing), driver and CUDA stack lifecycle, capacity planning, and throughput/latency/utilization tuning.
¢ Run the model registry with versioning, lineage, approval workflow and model-risk status for platform-owned models.
¢ Build LLMOps capabilities: prompt and context versioning, evaluation harnesses, regression testing and token-cost monitoring.
¢ Stand up LLM observability and evaluation with tools such as Langfuse, Arize Phoenix and RAGAS.
¢ Integrate guardrails enforcement so governance policy executes at runtime.
¢ Build MLOps automation for platform-owned predictive models: training pipelines, feature store integration, drift detection and retraining triggers on Databricks.
¢ Operate model serving and training/evaluation environments on the shared substrate, benchmarking self-hosted models against provider APIs for cost and performance.
¢ Produce reproducible evidence for independent model validation and mentor junior AI/ML Platform Engineers.
Skills and Background Youll Need
Education: B.Tech / B.E. in Software Engineering, Computer Science, Data Science or Information Systems, or related field is required. Masters degree is preferred.
Experience:
- Typically requires 10+ years of software or ML engineering experience, including significant production ML/AI platform work.
- Expert Python and strong engineering fundamentals: testing, CI/CD and containers.
- Hands-on experience with Anthropic and OpenAI APIs and multi-provider gateway/routing patterns (LiteLLM-style).
- Hands-on GPU inference serving experience with NVIDIA NIM and/or open-source frameworks such as vLLM, SGLang.
- Experience operating GPU infrastructure on Kubernetes: GPU Operator, device plugins, MIG or time-slicing, and utilization/throughput tuning.
- MLflow and model registry experience, plus model serving and monitoring in production.
- LLM evaluation and observability experience with Langfuse, Arize Phoenix, RAGAS or equivalent.
- Experience implementing guardrails or safety controls for generative systems.
- Feature store experience and drift monitoring for predictive models
Preferred Experience: ¢ Experience deploying on Grace Blackwell-class systems (GB200/GB300 NVL72), NVIDIA AI Enterprise, NVIDIA Dynamo for disaggregated serving, or low-precision formats such as NVFP4/FP8.
¢ Databricks ML stack depth (Model Serving, Mosaic AI tooling) and Unity Catalog model governance.
¢ Retrieval engineering with vector stores or GraphRAG (Neo4j).
¢ Exposure to NIST AI Risk Management Framework or AI TRiSM in a regulated industry.
¢ Fine-tuning open-source models; GPU FinOps and inference cost optimization.
¢ Valuable to have Databricks Machine Learning Professional or Azure AI Engineer certification.
Key Behaviors of a Successful Candidate:
¢ Adaptability Seeks information to understand the rationale for and importance of the change.
¢ Customer Focus Displays an interest in the customer by trying to understand their concerns and issues; draws on customer insight to help others best meet current and future customer needs.
¢ Improvement Mindset Demonstrates curiosity by asking questions regarding current approaches/methods and identifying potential changes.
📌 AI ML Platform Engineer (Bengaluru)
🏢 The Standard India
📍 Bengaluru