14 Aug
|
NAVA
|
Bengaluru
About Nava
Nava is building next-generation AI infrastructure and inference platforms designed to power the future of AI. We are looking for a Head of GPU Cluster Engineering to lead the architecture, deployment, and operations of large-scale GPU clusters.
This is a highly strategic leadership role responsible for defining the engineering vision, driving execution across multiple infrastructure domains, and ensuring our GPU platforms deliver world-class performance, scalability, and reliability.
What You'll Do
Platform Leadership
- Own end-to-end engineering outcomes, from reference architecture through production operations.
- Define the technical vision, architecture, and operating model for large-scale GPU infrastructure.
- Establish engineering standards, design principles, and operational excellence across the platform.
- Drive platform scalability, resiliency, and performance to support rapidly growing AI workloads.
Architecture & Technical Strategy
- Lead the design and evolution of GPU cluster architectures across compute, networking, storage, orchestration, and observability.
- Evaluate and drive technology decisions around GPUs, interconnects, storage, Kubernetes, scheduling, and AI infrastructure software.
- Review and approve architecture decisions while balancing performance, reliability, cost, and operational complexity.
- Drive infrastructure standardization and automation across deployments.
Cross-Functional Leadership
- Lead execution across multiple engineering pillars:
- GPU Compute
- High-Speed Networking
- Storage
- Platform Engineering
- Site Reliability Engineering (SRE)
- Infrastructure Automation
- Partner closely with Product, Supply Chain, Data Centre Operations, and Customer Success teams to ensure successful platform delivery.
- Act as the technical escalation point for critical engineering decisions.
Delivery & Operational Excellence
- Own engineering readiness gates from design through production deployment.
- Drive release planning, operational reviews, risk assessments, and post-incident analysis.
- Establish SLAs, SLOs, and operational metrics for platform health and reliability.
- Champion automation, observability, incident management, and continuous improvement across the engineering organization.
Team Leadership
- Build, mentor, and scale a high-performing engineering organization.
- Develop technical leaders across infrastructure disciplines.
- Foster a culture of engineering excellence, ownership, collaboration, and innovation.
- Support hiring and talent development for critical infrastructure roles.
What We're Looking For
- 12 years of experience in infrastructure engineering, distributed systems, cloud platforms, or AI infrastructure.
- Proven experience leading large-scale infrastructure or platform engineering teams.
- Deep understanding of GPU clusters, AI infrastructure, HPC, or large-scale distributed systems.
- Solid expertise across:
- GPU Compute (NVIDIA ecosystem preferred)
- High-performance networking (InfiniBand, RoCE, RDMA)
- Kubernetes and container orchestration
- Distributed storage systems
- Infrastructure automation and observability
- Linux systems and platform engineering
- Experience designing highly available, scalable production infrastructure.
- Strong architectural thinking with the ability to balance technical excellence and business priorities.
- Excellent stakeholder management and cross-functional leadership skills.
Nice to Have
- Experience building AI factories, GPU cloud platforms, or large-scale inference infrastructure.
- Familiarity with CUDA, NCCL, Slurm, Ray, Kubeflow, or similar AI infrastructure technologies.
- Experience working with hyperscalers, cloud providers, or AI-first technology companies.
- Exposure to multi-region or global infrastructure deployments.
Why Join Nava?
- Build one of the world's leading AI infrastructure platforms.
- Lead the engineering strategy behind next-generation GPU clusters and AI factories.
- Work alongside world-class engineers solving some of the most challenging infrastructure problems in AI.
- Shape the future of AI infrastructure from architecture through production at global scale.
Skills: leadership,automation,platforms,gpu,reliability,storage,architecture,cloud,infrastructure,cluster,drive
📌 Head of GPU Cluster Engineering (Bengaluru)
🏢 NAVA
📍 Bengaluru