Consultant - GPU Performance Engineer (Gurugram)

Consultant - GPU Performance Engineer (Gurugram)

10 Aug
|
Credflow AI
|
Gurugram

10 Aug

Credflow AI

Gurugram

Consulting Engineer – AI Inference Infrastructure (Remote)

Commitment: 20 hrs/week (Flexible)
Duration: 6–12 weeks (Renewable)
Location: Remote (Weekly overlap with US Pacific time required)

About PrimaLabs
PrimaLabs is building a workload-specific AI inference serving platform that maximizes real-world accelerator performance through closed-loop autotuning.
We're looking for experts in one or two of the following areas:
Disaggregated Inference
KV Cache Systems (LMCache, HiCache, etc.)
Speculative Decoding
Inference Routing & Gateway
Serving Stack Engineering (vLLM, SGLang, NVIDIA Dynamo)
Throughput & Latency Optimization

Requirements
5+ years in Systems, HPC, Distributed Systems, or ML Infrastructure
Hands-on experience with AI inference serving
Strong Python; working knowledge of C++/CUDA




Experience with multi-node GPU clusters and containerized deployments
Benchmark-driven, performance optimization mindset
Ability to work independently with weekly checkpoints

Nice to Have
Quantized inference (FP8/FP4)
KV transfer/cache-tiering projects
Open-source contributions (vLLM, SGLang, etc.)
Bayesian optimization or autotuning experience

What You'll Get
Access to production-grade multi-GPU clusters
Real-world traffic and benchmarking infrastructure
Direct collaboration with the founding team
Rapid decision-making and high-impact engineering work

📌 Consultant - GPU Performance Engineer (Gurugram)
🏢 Credflow AI
📍 Gurugram

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: consultant - gpu performance engineer (gurugram) / gurugram

Subscribe to this job alert:

Get the latest job offers by email for: consultant - gpu performance engineer (gurugram) / gurugram