14 Aug
|
NAVA
|
Bengaluru
About Nava
Nava is building next-generation AI infrastructure and inference platforms designed to power enterprise AI at scale. We're looking for a Head of Compute &
- Inference Platform to lead the engineering strategy and execution for our compute orchestration and inference platform.
This role is responsible for building the platform that efficiently shares GPU resources, serves AI models at scale, and delivers high-performance, multi-tenant AI infrastructure to customers. You'll own everything from workload scheduling and GPU allocation to model serving, runtime optimization, APIs, and platform performance.
What You'll Do
Compute &
- Inference Platform Strategy
- Define the technical vision and roadmap for Nava's Compute &
- Inference Platform.
- Own the architecture and evolution of GPU scheduling, workload orchestration, and inference serving infrastructure.
- Build a highly scalable, secure, and multi-tenant AI platform capable of serving enterprise workloads.
- Drive platform innovation while balancing performance, reliability, cost, and customer experience.
GPU Resource Management
- Design and optimize GPU scheduling, allocation, and capacity management across large-scale GPU clusters.
- Develop intelligent scheduling strategies to maximize GPU utilization while maintaining fairness and workload isolation.
- Own resource quotas, workload prioritization, tenancy management, and capacity planning.
- Continuously improve infrastructure efficiency and cost optimization.
Inference Platform &
- Model Serving
- Own the end-to-end model serving stack, including deployment, scaling, and lifecycle management.
- Lead engineering for model runtimes, inference frameworks, serving infrastructure, and API gateways.
- Ensure rapid onboarding and deployment of recent AI models while maintaining platform stability and performance.
- Optimize inference latency, throughput, and infrastructure utilization.
Platform APIs &
- Developer Experience
- Own customer-facing APIs, SDKs, and platform interfaces that enable seamless deployment and management of AI workloads.
- Improve developer experience through automation, self-service capabilities,
and platform tooling.
- Work closely with Product teams to define platform capabilities and customer-facing features.
Performance &
- Cost Optimization
- Establish platform performance benchmarks, service level objectives (SLOs), and cost optimization targets.
- Continuously improve GPU utilization, inference efficiency, scheduling algorithms, and workload performance.
- Drive benchmarking, performance testing, and capacity planning initiatives.
- Build observability and telemetry to measure platform health and customer experience.
Cross-Functional Leadership
- Collaborate closely with GPU Cluster Engineering, Platform Reliability, AI Infrastructure Security, Networking, Product, and Customer Success teams.
- Align infrastructure capabilities with product strategy and customer requirements.
- Act as the technical leader for platform architecture and major engineering decisions.
Team Leadership
- Build and lead a high-performing Compute &
- Inference engineering organization.
- Mentor engineering managers and senior engineers.
- Foster a culture of technical excellence, ownership, innovation, and operational discipline.
Success Metrics
Impact in this role will be measured by clear, outcome-driven milestones:
- Infrastructure Efficiency: Achieve and sustain ≥85% average GPU utilization across production clusters.
- Inference Performance: Deliver sub-50ms p99 latency for high-volume models, with ≥95% SLA adherence for throughput targets.
- Developer Velocity: Reduce time-to-deploy new models from days to minutes, supporting ≥200 model deployments/month.
- Reliability &
- Isolation: Maintain ≥99.95% platform uptime and enforce strict performance isolation across tenants (≤5% variance under load).
- Customer Experience:
Achieve ≥4.5/5 NPS on platform usability and ≥99.9% API availability.
- Cost Efficiency: Decrease cost per inference by ≥20% YoY while scaling throughput by ≥2x.
- Time-to-Value: Reduce time-to-production for new models by ≥50% over 12 months.
What We're Looking For
- 12+ years of experience building large-scale distributed systems, cloud platforms, AI infrastructure, or compute platforms.
- Proven experience leading platform engineering teams responsible for production-scale infrastructure.
- Deep expertise in:
- GPU scheduling and resource management
- Kubernetes and container orchestration
- Distributed systems and cloud-native platforms
- AI model serving and inference architectures
- Multi-tenant platform design
- API platforms and developer tooling
- Performance engineering and capacity planning
- Strong understanding of AI infrastructure technologies, including model serving frameworks and orchestration platforms.
- Excellent architectural thinking with the ability to balance scalability, performance, security, and operational simplicity.
- Exceptional leadership, stakeholder management, and communication skills.
Nice to Have
- Experience with NVIDIA GPU ecosystems, CUDA, Triton Inference Server, vLLM, Ray Serve, KServe, or similar inference technologies.
- Familiarity with LLM serving, distributed inference, model optimization, and quantization techniques.
- Experience building AI cloud platforms, GPU-as-a-Service offerings, or hyperscale infrastructure.
- Exposure to HPC environments and large-scale enterprise AI deployments.
Why Join Nava?
- Lead the engineering vision for one of the world's most advanced AI inference platforms.
- Solve some of the most challenging problems in GPU scheduling, distributed inference, and AI infrastructure.
- Work alongside world-class engineers building the future of enterprise AI.
- Shape the platform that powers next-generation AI applications at global scale.
Skills: infrastructure,scheduling,gpu,enterprise,orchestration,multi-tenant,management,building,optimization,inference,customer,platforms
📌 Head of Compute & Inference Platform (Bengaluru)
🏢 NAVA
📍 Bengaluru