11 Sep
|
Data Economy
|
Hyderabad
11 Sep
Data Economy
Hyderabad
Job Summary
Job Title: MLOps / Serving Engineer
Experience: 5+ years
Location: Hyderabad or Pune
Notice Period: 0-30 days
Work mode: Hybrid
Responsibilities
- Deploy fine-tuned LLMs using vLLM, TensorRT-LLM, or Triton with continuous batching on AWS GPU instances
- Build shadow-mode deployment: run fine-tuned model alongside production, log comparison data without impacting live traffic
- Execute staged rollout: canary (5%) gradual ramp (25% 50% 100%) with automated rollback on quality degradation
- Optimize inference for input-heavy workloads (~17K token inputs, ~130 token outputs): prefill throughput, KV-cache, INT8 quantization
- Build monitoring dashboards: latency, throughput, accuracy metrics, cost per request
- Design auto-scaling; implement high-availability (2 instances); automated rollback triggers on end-to-end quality metrics
Requirements
- 5+ years MLOps or ML infrastructure engineering
- Hands-on with vLLM, TensorRT-LLM, or Triton Inference Server
- Deep familiarity with g5, p4de, p5 instance families, EC2 auto-scaling
- Have worked on Deployment patterns like Shadow-mode, canary, A/B traffic routing, automated rollback
- Experience onto Continuous batching, INT8 quantization, KV-cache management
- Expertise on Docker, Kubernetes (EKS) for ML workloads
- Worked on CloudWatch, Prometheus, Grafana
Job Type
Full time Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 MLOps / Serving Engineer (Hyderabad)
🏢 Data Economy
📍 Hyderabad