09 Oct
|
Rafay Systems
|
India
09 Oct
Rafay Systems
India
Role Overview
This engineering role begins where model training ends. Once a machine learning or large language model (LLM) is trained, an inference optimization engineer figures out how to serve it efficiently to real-
world users under strict speed and cost limits. The position sits between research and low-level systems engineering. It requires balancing speed—measured as Time-to-First-Token (TTFT) and tokens per second—with output accuracy and hardware costs.
Key Responsibilities
- Model Compression: Apply techniques like quantization (FP8, INT8, INT4), pruning, and distillation to shrink model size without losing accuracy.
- Runtime & Serving Tuning: Configure and optimize high-throughput serving frameworks such as vLLM, SGLang, TensorRT-LLM, or Triton Inference Server.
- Memory & Caching Management: Optimize KV cache usage, continuous batching, and speculative decoding to handle long context windows and multiple concurrent users.
- Profiling & Benchmarking: Use tools like Nsight Systems,
PyTorch Profiler, and Triton Metrics to identify memory bandwidth or compute bottlenecks on GPU/TPU clusters.
- Hardware Co-Optimization: Tune custom kernels (CUDA/Triton) and manage distributed setups across multi-GPU clusters using tensor and pipeline parallelism
Key Requirements & Qualifications
- Experience: 4+ years in systems programming, high-performance computing (HPC), or ML infrastructure.
- Core Languages: Strong proficiency in Python or Go, with familiarity reading or writing CUDA/Triton kernels.
- Frameworks: Hands-on production experience with up-to-date inference engines (vLLM, TensorRT-LLM, SGLang) and profiling tools (Nsight).
- Foundational Knowledge: Deep understanding of transformer architectures, memory hierarchies, arithmetic intensity, and roofline performance models.
📌 Inference Optimization Engineer (Sr Engineer / Sta5 Engineer / Principal Engineer) (India)
🏢 Rafay Systems
📍 India