Inference Optimization Engineer (Sr Engineer / Sta5 Engineer / Principal Engineer) (India)

Inference Optimization Engineer (Sr Engineer / Sta5 Engineer / Principal Engineer) (India)

09 Oct
|
Rafay Systems
|
India

09 Oct

Rafay Systems

India

Role Overview

This engineering role begins where model training ends. Once a machine learning or large language model (LLM) is trained, an inference optimization engineer figures out how to serve it efficiently to real-

world users under strict speed and cost limits. The position sits between research and low-level systems engineering. It requires balancing speed—measured as Time-to-First-Token (TTFT) and tokens per second—with output accuracy and hardware costs.

Key Responsibilities

- Model Compression: Apply techniques like quantization (FP8, INT8, INT4), pruning, and distillation to shrink model size without losing accuracy.

- Runtime & Serving Tuning: Configure and optimize high-throughput serving frameworks such as vLLM, SGLang, TensorRT-LLM, or Triton Inference Server.

- Memory & Caching Management: Optimize KV cache usage, continuous batching, and speculative decoding to handle long context windows and multiple concurrent users.

- Profiling & Benchmarking: Use tools like Nsight Systems,



PyTorch Profiler, and Triton Metrics to identify memory bandwidth or compute bottlenecks on GPU/TPU clusters.

- Hardware Co-Optimization: Tune custom kernels (CUDA/Triton) and manage distributed setups across multi-GPU clusters using tensor and pipeline parallelism

Key Requirements & Qualifications

- Experience: 4+ years in systems programming, high-performance computing (HPC), or ML infrastructure.

- Core Languages: Strong proficiency in Python or Go, with familiarity reading or writing CUDA/Triton kernels.

- Frameworks: Hands-on production experience with up-to-date inference engines (vLLM, TensorRT-LLM, SGLang) and profiling tools (Nsight).

- Foundational Knowledge: Deep understanding of transformer architectures, memory hierarchies, arithmetic intensity, and roofline performance models.

📌 Inference Optimization Engineer (Sr Engineer / Sta5 Engineer / Principal Engineer) (India)
🏢 Rafay Systems
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: inference optimization engineer (sr engineer / sta5 engineer / principal engineer) (india) / india

Subscribe to this job alert:

Get the latest job offers by email for: inference optimization engineer (sr engineer / sta5 engineer / principal engineer) (india) / india