We serve inference at $/token margins that dont tolerate sloppy stacks. Youll own the serving layer vLLM, TensorRT-LLM, Triton and the benchmarking discipline that keeps it honest.
Youll own
- Serving-stack selection per workload (continuous batching vs. static, KV cache strategy, paged attention).
- Quantization (FP8, AWQ, GPTQ) and the eval harness that proves the trade-offs.
- The InferenceBench-style benchmarks that compare our serving against the field.
You probably have
- Shipped at least one production inference stack on H100s or A100s.
- Read a recent vLLM PR for fun.
- Solid opinions about speculative decoding.
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 ML Infrastructure Engineer (India)
🏢 Yobitel Communications
📍 India
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.