08 Oct
|
Gnani Innovations Private
|
Bengaluru
08 Oct
Gnani Innovations Private
Bengaluru
Role Overview:We are hiring an LLMOps Engineer to own how our large language models actually run in production - the serving stack, the optimization work that makes it affordable, and the low-level performance engineering that makes it fast. This is a deep infrastructure and performance role, not a wrapper-and-API role.You will work on self-hosted foundational models served on large-scale NVIDIA GPU clusters, under real-time latency budgets set by live spoken conversations. Voice inference is unforgiving: time-to-first-token is a product feature, not a number on a dashboard, and every millisecond of tail latency is audible to a caller. Your mandate is to hold latency and quality while driving cost per million tokens down.The work spans three layers: the serving engine (vLLM, SGLang, NVIDIA Dynamo), the model itself (quantization, distillation, speculative decoding, pruning), and the kernels underneath (CUDA/Triton, attention and MoE kernels, profiling and bottleneck analysis). We expect real depth in at least two of the three and genuine curiosity about the third.Key Responsibilities:Inference Serving and Platform:- Own the production LLM serving stack across vLLM, SGLang, and NVIDIA Dynamo - including engine selection per workload, with benchmark evidence to justify it- Tune the serving path end to end: continuous batching, chunked prefill, paged and radix attention, prefix and KV caching, cache-aware and sticky routing, speculative decode integration- Design and operate disaggregated prefill/decode and multi-node deployments; configure tensor, pipeline, and expert parallelism for both dense and Mixture-of-Experts models- Run inference on Kubernetes with autoscaling, rolling model updates, canary releases, and tested rollback paths; enforce staging-to-production promotion gates- Instrument everything that matters: TTFT, inter-token latency, p50/p95/p99, tokens/sec/GPU, KV cache hit rate, queue depth, GPU utilization and MFU, and cost per million tokens- Build and maintain internal serving APIs, model registries, and deployment tooling so that model updates are routine rather than eventsModel Optimization:- Quantization: FP8, INT8, and INT4 weight and activation quantization plus KV cache quantization - calibration set design, per-layer sensitivity analysis, and accuracy recovery, with measured deltas per language and task- Speculative decoding: draft-model, EAGLE/Medusa-style, and n-gram approaches; tuned for real acceptance rate and end-to-end latency gain under production traffic, not theoretical speedup- Distillation:
build smaller task-specific students from larger teachers for latency-critical paths, held to explicit eval-parity targets- Pruning and adapters: structured pruning, LoRA/adapter serving, and multi-adapter batching for per-deployment specialization- Compilation: torch.compile, CUDA graphs, and TensorRT-LLM engine builds - including diagnosing recompilation and dynamic-shape stallsKernel and Low-Level Performance:- Profile with Nsight Systems/Compute and the PyTorch profiler; classify bottlenecks as memory-bound, compute-bound, or launch/synchronization-bound, and act on the classification- Write and tune custom kernels in CUDA and Triton - attention variants, fused MoE dispatch, sampling, quantized GEMM- Eliminate host-device synchronization stalls, dynamic-shape recompilation, and unoverlapped collectives in distributed serving- Work fluently with attention and MoE kernel libraries (FlashAttention, FlashInfer, CUTLASS/cuBLAS, expert-parallel dispatch libraries) and know when to use them versus write your own- Reason from first principles about arithmetic intensity, memory bandwidth, and roofline limits before reaching for a toolEvaluation, Reliability, and Cost:- Own the optimization regression gate: no optimized build reaches production without an accuracy and behavior evaluation across languages and task types- Build load-testing harnesses that replay realistic traffic - concurrency, length distributions, burstiness, multi-turn sessions- Run capacity planning and cost modeling; set and hit targets for cost per million tokens and per concurrent session- Carry on-call for inference services, write runbooks, and lead blameless postmortems on latency and availability incidents- Work closely with the training and post-training teams so that serving constraints inform model architecture decisions early, not after the checkpoint landsMust Have:- 6 - 10 years total experience, with 3+ years owning large-scale LLM or speech inference in production- Has owned an inference platform end to end: architecture, SLOs, capacity, cost, deployment safety, and on-call- Has written or substantially tuned custom CUDA or Triton kernels that shipped to production- MoE serving at scale:
expert parallelism, routing load imbalance, dispatch kernels, and the failure modes specific to sparse models- Deep profiling ability - has diagnosed and fixed a utilization or MFU collapse in a distributed multi-node system, and can walk through the investigation- Track record running an accuracy-preserving optimization program (quantization plus speculative decoding plus distillation) behind real evaluation gates- Sets technical direction, mentors engineers, and can make a defensible build-versus-adopt call on serving infrastructure- Upstream contributions to open-source serving or kernel projects (vLLM, SGLang, FlashInfer, TensorRT-LLM, or similar) - a strong plus, and close to an expectation at this levelGood to Have:- Real-time speech inference: streaming ASR/TTS, sub-second time-to-first-audio budgets, barge-in and turn-taking constraints- Serving hybrid-architecture models (state-space/Mamba blocks combined with attention) and understanding their distinct state and cache management- TensorRT-LLM engine building, NVIDIA NIM, or Triton Inference Server in production- Experience with Indic or other multilingual and code-mixed models, and the evaluation discipline that comes with them- Ray Serve, KServe, LLM gateways, or semantic and prefix cache layers- Enough training-side exposure (Megatron, NeMo, DeepSpeed, FSDP) to work fluently with the pretraining team- Public benchmarks, blog posts, or talks on inference optimizationWhat You Will Work On:You will work on in-house foundational models - not third-party API endpoints - running on large NVIDIA GPU clusters and serving live enterprise and government-scale voice traffic. The models are ours, so the whole stack is open to you: you can change the serving engine, the quantization recipe, the kernel, or the model itself, and you will often need to change more than one to hit a target.The constraints are real and the feedback loop is fast. A latency regression is heard by callers within minutes. A successful optimization shows up directly in infrastructure spend. You will have access to proprietary speech and language data, in-house training and evaluation infrastructure, a close engineering partnership with NVIDIA, and a team that has been building Indian-language AI since 2016.This role has a transparent path to owning the inference platform and its technical direction, or to a deeper specialization in performance and kernel engineering, depending on where your strength lies. (ref:hirist.tech)
📌 Gnani.ai - LLMOps Engineer (Bengaluru)
🏢 Gnani Innovations Private
📍 Bengaluru