Inference Systems Engineer (India)

Inference Systems Engineer (India)

28 Aug
|
LambdaQ labs
|
India

28 Aug

LambdaQ labs

India

About The Role

You'll build and optimize the serving stack that makes Vikasit Inference rapid, reliable, and affordable at India scale — kernels, batching, caching, and the OpenAI-compatible API surface developers love.

What you'll do

- —Optimize throughput and latency across our model fleet (vLLM / SGLang / TensorRT-LLM)
- —Build batching, KV-cache, quantization, and routing for MoE models
- —Own reliability, autoscaling, and cost of the inference platform
- —Keep the OpenAI-compatible API rock-solid for streaming and tool calls

What we're looking for

- —Deep systems engineering (GPU, CUDA-adjacent, or high-performance serving)
- —Experience with an LLM serving framework in production
- —Comfort owning latency, throughput, and cost SLOs

Nice to have

- —CUDA / Triton kernels
- —Quantization (GGUF, AWQ, FP8)
- —Multi-region infra

Sound like you?

We hire for skill over credentials. Tell us why you're a fit — links and projects welcome.

Apply for this role

📌 Inference Systems Engineer (India)
🏢 LambdaQ labs
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: inference systems engineer (india) / india

Subscribe to this job alert:

Get the latest job offers by email for: inference systems engineer (india) / india