Consulting Engineer – AI Inference Infrastructure (Remote)
Commitment: 20 hrs/week (Adaptable)
Duration: 6–12 weeks (Renewable)
Location: Remote (Weekly overlap with US Pacific time required)
About PrimaLabs
PrimaLabs is building a workload-specific AI inference serving platform that maximizes real-world accelerator performance through closed-loop autotuning.
We're looking for experts in one or two of the following areas:
Disaggregated Inference
KV Cache Systems (LMCache, HiCache, etc.)
Speculative Decoding
Inference Routing & Gateway
Serving Stack Engineering (vLLM, SGLang, NVIDIA Dynamo)
Throughput & Latency Optimization
Requirements
5+ years in Systems, HPC, Distributed Systems, or ML Infrastructure
Hands-on experience with AI inference serving
Strong Python; working knowledge of C++/CUDA
Experience with multi-node GPU clusters and containerized deployments
Benchmark-driven, performance optimization mindset
Ability to work independently with weekly checkpoints
Nice to Have
Quantized inference (FP8/FP4)
KV transfer/cache-tiering projects
Open-source contributions (vLLM, SGLang, etc.)
Bayesian optimization or autotuning experience
What You'll Get
Access to production-grade multi-GPU clusters
Real-world traffic and benchmarking infrastructure
Direct collaboration with the founding team
Fast decision-making and high-impact engineering work