06 Aug
|
NAVA
|
Bengaluru
Role & Responsibilities
Design and deploy low-latency inference pipelines for LLMs, diffusion models, and vision transformers across CPU, GPU, and NPU architectures.
Optimize model serving stacks using Triton Inference Server, vLLM, or TGI—tuning batch sizes, quantization, and memory layout for peak performance.
Containerize and orchestrate inference services via Docker and Kubernetes, ensuring high availability and auto-scaling under fluctuating workloads.
Implement model monitoring, health checks, and A/B testing frameworks to validate performance and drift in production.
Collaborate with ML Engineers to convert trained models into production-ready formats (ONNX, TensorRT, GGUF) with minimal accuracy loss.
Build observability dashboards (Prometheus/Grafana)
and alerting logic to detect and mitigate inference bottlenecks in real time.
Skills & Qualifications
Must-Have
Python
Docker
Kubernetes
Triton Inference Server
ONNX Runtime
PyTorch
TensorRT
Prometheus
Grafana
CI/CD (GitHub Actions, GitLab CI)
Preferred
Experience with vLLM or TGI
Knowledge of Model Quantization (AWQ, GPTQ, GGUF)
Familiarity with NVIDIA Triton model ensemble pipelines
Skills: cuda,architecture,building,foundation,ml,management,models,distributed systems,infrastructure,inference
📌 Inference Systems Engineer (Bengaluru)
🏢 NAVA
📍 Bengaluru