09 Sep
|
gnani.ai
|
Bengaluru
09 Sep
gnani.ai
Bengaluru
About Gnani.ai
Gnani.ai is India's leading enterprise voice AI company, building the language infrastructure that powers intelligent voice agents, speech recognition, and text-to-speech systems at scale across 40+ languages. Our models are deployed across government, BFSI, telecom, and enterprise verticals, processing over 30 million voice AI calls a day. Founded in 2016 and backed by Samsung Ventures and Info Edge Ventures, we are one of four companies selected under the IndiaAI Mission to build foundational AI models for India.
Role Overview
We are hiring an LLMOps Engineer to own how our large language models actually run in production — the serving stack, the optimization work that makes it affordable, and the low-level performance engineering that makes it rapid. This is a deep infrastructure and performance role, not a wrapper-and-API role.
You will work on self-hosted foundational models served on large-scale NVIDIA GPU clusters, under real-time latency budgets set by live spoken conversations. Voice inference is unforgiving: time-to-first-token is a product feature,
not a number on a dashboard, and every millisecond of tail latency is audible to a caller. Your mandate is to hold latency and quality while driving cost per million tokens down.
The work spans three layers: the serving engine (vLLM, SGLang, NVIDIA Dynamo), the model itself (quantization, distillation, speculative decoding, pruning), and the kernels underneath (CUDA/Triton, attention and MoE kernels, profiling and bottleneck analysis). We expect real depth in at least two of the three and genuine curiosity about the third.
Key Responsibilities :
Inference Serving and Platform
• Own the production LLM serving stack across vLLM, SGLang, and NVIDIA Dynamo — including engine selection per workload, with benchmark evidence to justify it
• Tune the serving path end to end: continuous batching, chunked prefill, paged and radix attention, prefix and KV caching, cache-aware and sticky routing, speculativ
📌 LLM Ops Engineer (Bengaluru)
🏢 gnani.ai
📍 Bengaluru