06 Aug
|
NAVA
|
Bengaluru
Role & Responsibilities
Optimize LLM inference pipelines for latency, throughput, and memory efficiency across GPU/TPU hardware—using quantization, pruning, kernel fusion, and runtime scheduling.
Implement and benchmark model compression techniques (INT4/INT8/FP8, LoRA, Sparse attention) to reduce model size while preserving performance.
Integrate and tune LLM serving frameworks (vLLM, TensorRT-LLM, HuggingFace TGI, ONNX Runtime) for high-throughput, low-latency inference in cloud and on-premise environments.
Collaborate with ML researchers and infrastructure engineers to profile bottlenecks and implement hardware-aware optimizations using CUDA, Triton, or custom kernels.
Design and automate model benchmarking suites to track performance regression, memory footprint, and cost-per-token across model versions and hardware targets.
Document optimization playbooks and contribute to internal tooling for repeatable, scalable model deployment workflows.
Skills & Qualifications
Must-Have
PyTorch
TensorRT
vLLM
Quantization (INT4/INT8/FP8)
CUDA
ONNX Runtime
Triton Inference Server
LLM Inference Optimization
Preferred
Experience with Mixture-of-Experts (MoE) models
Familiarity with HuggingFace Transformers and TGI
Knowledge of NPU/ASIC inference backends (e.g., Qualcomm, Groq, Cerebras)
Skills: cuda,ml,learning,compression,optimization,distillation,decoding,research,machine learning,training
📌 Post-Training Optimization Engineer (Bengaluru)
🏢 NAVA
📍 Bengaluru