13 Aug
|
NAVA
|
Bengaluru
Role & Responsibilities
- Optimize LLM inference pipelines for latency, throughput, and memory efficiency across GPU/TPU hardware—using quantization, pruning, kernel fusion, and runtime scheduling.
- Implement and benchmark model compression techniques (INT4/INT8/FP8, LoRA, Sparse attention) to reduce model size while preserving performance.
- Integrate and tune LLM serving frameworks (vLLM, TensorRT-LLM, HuggingFace TGI, ONNX Runtime) for high-throughput, low-latency inference in cloud and on-premise environments.
- Collaborate with ML researchers and infrastructure engineers to profile bottlenecks and implement hardware-aware optimizations using CUDA, Triton, or custom kernels.
- Design and automate model benchmarking suites to track performance regression, memory footprint, and cost-per-token across model versions and hardware targets.
- Document optimization playbooks and contribute to internal tooling for repeatable, scalable model deployment workflows.
Skills & Qualifications
Must-Have
- PyTorch
- TensorRT
- vLLM
- Quantization (INT4/INT8/FP8)
- CUDA
- ONNX Runtime
- Triton Inference Server
- LLM Inference Optimization
Preferred
- Experience with Mixture-of-Experts (MoE) models
- Familiarity with HuggingFace Transformers and TGI
- Knowledge of NPU/ASIC inference backends (e.g., Qualcomm, Groq, Cerebras)
Skills: cuda,ml,learning,compression,optimization,distillation,decoding,research,machine learning,training
📌 Post-Training Optimization Engineer (Bengaluru)
🏢 NAVA
📍 Bengaluru