09 Sep
|
Gnani Innovations
|
Bengaluru
09 Sep
Gnani Innovations
Bengaluru
Bengaluru
6 - 10 years experience
Responsibilities
About Gnani.ai Gnani.ai is India's leading enterprise voice AI company, building the language infrastructure that powers intelligent voice agents, speech recognition, and text-to-speech systems at scale across 40+ languages. Our models are deployed across government, BFSI, telecom, and enterprise verticals, processing over 30 million voice AI calls a day. Founded in 2016 and backed by Samsung Ventures and Info Edge Ventures, we are one of four companies selected under the IndiaAI Mission to build foundational AI models for India.
Role Overview
We are hiring an LLMOps Engineer to own how our large language models actually run in production — the serving stack, the optimization work that makes it affordable, and the low-level performance engineering that makes it fast. This is a deep infrastructure and performance role, not a wrapper-and-API role.
You will work on self-hosted foundational models served on large-scale NVIDIA GPU clusters, under real-time latency budgets set by live spoken conversations. Voice inference is unforgiving: time-to-first-token is a product feature, not a number on a dashboard, and every millisecond of tail latency is audible to a caller. Your mandate is to hold latency and quality while driving cost per million tokens down.
The work spans three layers: the serving engine (vLLM, SGLang, NVIDIA Dynamo), the model itself (quantization, distillation, speculative decoding, pruning), and the kernels underneath (CUDA/Triton, attention and MoE kernels, profiling and bottleneck analysis). We expect real depth in at least two of the three and genuine curiosity about the third.
Key Responsibilities
Inference Serving and Platform
- Own the production LLM serving stack across vLLM, SGLang, and NVIDIA Dynamo — including engine selection per workload, with benchmark evidence to justify it
- Tune the serving path end to end: continuous batching, chunked prefill, paged and radix attention, prefix and KV caching, cache-aware and sticky routing, speculative decode integration
- Design and operate disaggregated prefill/decode and multi-node deployments; configure tensor, pipeline, and expert parallelism for both dense and Mixture-of-Experts models
- Run inference on Kubernetes with autoscaling, rolling model updates, canary releases, and tested rollback paths; enforce staging-to-production promotion gates
- Instrument everything that matters: TTFT, inter-token latency, p50/p95/p99, tokens/sec/GPU, KV cache hit rate, queue depth, GPU utilization and MFU, and cost per million tokens
- Build and maintain internal serving APIs, model registries, and deployment tooling so that model updates are routine rather than events
Model Optimization
- Quantization: FP8, INT8,
and INT4 weight and activation quantization plus KV cache quantization — calibration set design, per-layer sensitivity analysis, and accuracy recovery, with measured deltas per language and task
- Speculative decoding: draft-model, EAGLE/Medusa-style, and n-gram approaches; tuned for real acceptance rate and end-to-end latency gain under production traffic, not theoretical speedup
- Distillation: build smaller task-specific students from larger teachers for latency-critical paths, held to explicit eval-parity targets
- Pruning and adapters: structured pruning, LoRA/adapter serving, and multi-adapter batching for per-deployment specialization
- Compilation: torch.compile, CUDA graphs, and TensorRT-LLM engine builds — including diagnosing recompilation and agile-shape stalls
Kernel and Low-Level Performance
- Profile with Nsight Systems/Compute and the PyTorch profiler; classify bottlenecks as memory-bound, compute-bound, or launch/synchronization-bound, and act on the classification
- Write and tune custom kernels in CUDA and Triton — attention variants, fused MoE dispatch, sampling, quantized GEMM
- Eliminate host-device synchronization stalls, dynamic-shape recompilation, and unoverlapped collectives in distributed serving
- Work fluently with attention and MoE kernel libraries (FlashAttention, FlashInfer, CUTLASS/cuBLAS, expert-parallel dispatch libraries) and know when to use them versus write your own
- Reason from first principles about arithmetic intensity, memory bandwidth, and roofline limits before reaching for a tool
Evaluation, Reliability, and Cost
- Own the optimization regression gate: no optimized build reaches production without an accuracy and behavior evaluation across languages and task types
- Build load-testing harnesses that replay realistic traffic — concurrency, length distributions, burstiness, multi-turn sessions
- Run capacity planning and cost modeling; set and hit targets for cost per million tokens and per concurrent session
- Carry on-call for inference services, write runbooks, and lead blameless postmortems on latency and availability incidents
- Work closely with the training and post-training teams so that serving constraints inform model architecture decisions early, not after the checkpoint lands
Must Have
- 6 to 10 years total experience, with 3+ years owning large-scale LLM or speech inference in production
- Has owned an inference platform end to end: architecture, SLOs, capacity, cost,
deployment safety, and on-call
- Has written or substantially tuned custom CUDA or Triton kernels that shipped to production
- MoE serving at scale: expert parallelism, routing load imbalance, dispatch kernels, and the failure modes specific to sparse models
- Deep profiling ability — has diagnosed and fixed a utilization or MFU collapse in a distributed multi-node system, and can walk through the investigation
- Track record running an accuracy-preserving optimization program (quantization plus speculative decoding plus distillation) behind real evaluation gates
- Sets technical direction, mentors engineers, and can make a defensible build-versus-adopt call on serving infrastructure
- Upstream contributions to open-source serving or kernel projects (vLLM, SGLang, FlashInfer, TensorRT-LLM, or similar) — a strong plus, and close to an expectation at this level
Good to Have
- Real-time speech inference: streaming ASR/TTS, sub-second time-to-first-audio budgets, barge-in and turn-taking constraints
- Serving hybrid-architecture models (state-space/Mamba blocks combined with attention) and understanding their distinct state and cache management
- TensorRT-LLM engine building, NVIDIA NIM, or Triton Inference Server in production
- Experience with Indic or other multilingual and code-mixed models, and the evaluation discipline that comes with them
- Ray Serve, KServe, LLM gateways, or semantic and prefix cache layers
- Enough training-side exposure (Megatron, NeMo, DeepSpeed, FSDP) to work fluently with the pretraining team
- Public benchmarks, blog posts, or talks on inference optimization
What You Will Work On You will work on in-house foundational models — not third-party API endpoints — running on large NVIDIA GPU clusters and serving live enterprise and government-scale voice traffic. The models are ours, so the whole stack is open to you: you can change the serving engine, the quantization recipe, the kernel, or the model itself, and you will often need to change more than one to hit a target.
The constraints are real and the feedback loop is fast. A latency regression is heard by callers within minutes. A successful optimization shows up directly in infrastructure spend. You will have access to proprietary speech and language data, in-house training and evaluation infrastructure, a close engineering partnership with NVIDIA, and a team that has been building Indian-language AI since 2016.
This role has a clear path to owning the inference platform and its technical direction, or to a deeper specialization in performance and kernel engineering, depending on where your strength lies.
Skills Required
Primary Skills RAG, Agentic Systems, vLLM/SGLang & Vector Databases
Agentic & LLM
TensorRT-LLM
SLOs
📌 Sr LLMOps Engineer -— LLM Inference & Model Optimization (Bengaluru)
🏢 Gnani Innovations
📍 Bengaluru