07 Aug
|
USHAK KAAL
|
Gurugram
07 Aug
USHAK KAAL
Gurugram
Consulting Engineer – AI Inference Infrastructure (Remote) Commitment: 20 hrs/week (Adaptable)
Duration: 6–12 weeks (Renewable)
Location: Remote (Weekly overlap with US Pacific time required) About PrimaLabs
PrimaLabs is building a workload-specific AI inference serving platform that maximizes real-world accelerator performance through closed-loop autotuning.
We're looking for experts in one or two of the following areas:
Disaggregated Inference
KV Cache Systems (LMCache, HiCache, etc.)
Speculative Decoding
Inference Routing & Gateway
Serving Stack Engineering (vLLM, SGLang, NVIDIA Dynamo)
Throughput & Latency Optimization Requirements
5+ years in Systems, HPC, Distributed Systems, or ML Infrastructure
Hands-on experience with AI inference serving
Solid Python; working knowledge of C++/CUDA
Experience with multi-node GPU clusters and containerized deployments
Benchmark-driven, performance optimization mindset
Ability to work independently with weekly checkpoints Nice to Have
Quantized inference (FP8/FP4)
KV transfer/cache-tiering projects
Open-source contributions (vLLM, SGLang, etc.)
Bayesian optimization or autotuning experience What You'll Get
Access to production-grade multi-GPU clusters
Real-world traffic and benchmarking infrastructure
Direct collaboration with the founding team
Rapid decision-making and high-impact engineering work
📌 Consultant Gpu Performance Engineer Gurugram
🏢 USHAK KAAL
📍 Gurugram