14 Aug
|
NAVA
|
Bengaluru
About Nava Nava is building next-generation AI infrastructure and inference platforms that power enterprise AI at scale. We are looking for a Principal Engineer – GPU Orchestration to lead the design and evolution of our GPU orchestration platform, enabling efficient scheduling, resource sharing, and multi-tenant inference workloads. This role will own the orchestration layer that manages GPU resources for inference and small GPU clusters (2–4 node deployments), ensuring optimal utilization, rapid scaling, workload isolation, and an exceptional customer experience.
What You'll Do GPU Orchestration & Scheduling Own the architecture and implementation of GPU scheduling for inference workloads and small GPU clusters (2–4 node deployments).
Design intelligent workload scheduling algorithms that maximize GPU utilization while ensuring fairness and predictable performance.
Manage GPU allocation, quotas, and resource sharing across multiple customers and workloads.
Continuously optimize scheduling policies to improve efficiency and reduce infrastructure costs. Kubernetes & Platform Engineering Own the Kubernetes orchestration layer for AI workloads.
Design and maintain GPU-aware scheduling capabilities within Kubernetes.
Improve cluster lifecycle management, workload placement, and infrastructure automation.
Partner with Platform Engineering teams to continuously enhance cluster reliability and scalability. Multi-Tenancy & Resource Isolation Design secure multi-tenant GPU environments with solid workload isolation.
Implement resource quotas, admission controls, namespace isolation, and scheduling policies.
Ensure consistent customer experience while maintaining high infrastructure utilization.
Collaborate closely with the Security team on tenancy and access control mechanisms.
Autoscaling & Workload Optimization Build fast autoscaling capabilities for inference workloads.
Reduce cold-start times through intelligent provisioning and pre-warming strategies.
Optimize model placement based on workload characteristics, GPU availability, and latency requirements.
Improve workload elasticity while balancing cost and performance.
Performance Engineering
Define and monitor platform KPIs including:
GPU utilization
Scheduling latency
Cluster efficiency
Autoscaling performance
Cold-start latency
Workload throughput
Drive continuous improvements through benchmarking, performance tuning, and automation.
Cross-Functional Collaboration Partner with Compute & Inference Platform, GPU Cluster Engineering, Platform Reliability, AI Infrastructure Security, and Product teams.
Support onboarding of new AI models and customer workloads.
Contribute to platform architecture decisions across Nava's AI infrastructure.
Technical Leadership
Serve as the technical authority for GPU orchestration and workload scheduling.
Mentor senior engineers and contribute to engineering best practices.
Drive architectural reviews, technical design discussions, and long-term platform strategy.
Success Metrics You Will Be Measured
On GPU utilization across the platform
Scheduling efficiency and fairness
Autoscaling responsiveness
Cold-start latency
Multi-tenant performance and isolation
Platform reliability and scalability
Customer workload performance
Infrastructure cost optimization Qualifications Required Qualifications 10+ years of experience in distributed systems, cloud infrastructure, Kubernetes, or platform engineering.
Deep expertise in Kubernetes internals, scheduling, and container orchestration.
Strong understanding of GPU resource management and AI infrastructure.
Experience building large-scale scheduling and orchestration systems.
Expertise in
Kubernetes
Container runtimes
Distributed systems
Infrastructure automation
Resource scheduling
Multi-tenant platform architecture
Performance optimization
Strong software engineering skills in Go, Python, C++, or similar systems programming languages.
Excellent problem-solving, architecture, and technical leadership skills.
Preferred Qualifications Experience with NVIDIA GPUs, CUDA, MIG (Multi-Instance GPU), GPU Operator, or Kubernetes device plugins.
Familiarity with KServe, Ray Serve, Triton Inference Server, vLLM, Slurm, or similar AI infrastructure technologies.
Experience with large-scale inference platforms, GPU cloud providers, or HPC environments.
Knowledge of Kubernetes scheduler extensions, custom controllers, and operator development.
Why Join
Nava? Build the orchestration platform powering one of the world's leading AI inference platforms.
Solve challenging problems in GPU scheduling, resource optimization, and distributed systems.
Work with world-class engineers building cutting-edge AI infrastructure.
Shape the future of enterprise AI by enabling efficient, scalable, and secure GPU orchestration at global scale.
Skills: customer,architecture,infrastructure,gpu,design,utilization,kubernetes,cluster,orchestration,scheduling,isolation
📌 Principal Engineer – GPU Orchestration (Bengaluru)
🏢 NAVA
📍 Bengaluru