14 Aug
|
NAVA
|
Bengaluru
About Nava Nava is building next-generation AI infrastructure and inference platforms designed to power the future of AI. We are looking for a Head of GPU Cluster Engineering to lead the architecture, deployment, and operations of large-scale GPU clusters. This is a highly strategic leadership role responsible for defining the engineering vision, driving execution across multiple infrastructure domains, and ensuring our GPU platforms deliver world-class performance, scalability, and reliability.
What You'll Do Platform Leadership Own end-to-end engineering outcomes, from reference architecture through production operations.
Define the technical vision, architecture, and operating model for large-scale GPU infrastructure.
Establish engineering standards, design principles, and operational excellence across the platform.
Drive platform scalability, resiliency, and performance to support rapidly growing AI workloads. Architecture & Technical Strategy Lead the design and evolution of GPU cluster architectures across compute, networking, storage, orchestration, and observability.
Evaluate and drive technology decisions around GPUs, interconnects, storage, Kubernetes, scheduling, and AI infrastructure software.
Review and approve architecture decisions while balancing performance, reliability, cost, and operational complexity.
Drive infrastructure standardization and automation across deployments. Cross-Functional Leadership Lead execution across multiple engineering pillars:
GPU Compute
High-Speed Networking
Storage
Platform Engineering
Site Reliability Engineering (SRE)
Infrastructure Automation
Partner closely with Product, Supply Chain, Data Centre Operations, and Customer Success teams to ensure successful platform delivery.
Act as the technical escalation point for critical engineering decisions.
Delivery & Operational Excellence Own engineering readiness gates from design through production deployment.
Drive release planning, operational reviews, risk assessments, and post-incident analysis.
Establish SLAs, SLOs,
and operational metrics for platform health and reliability.
Champion automation, observability, incident management, and continuous improvement across the engineering organization.
Team Leadership
Build, mentor, and scale a high-performing engineering organization.
Develop technical leaders across infrastructure disciplines.
Foster a culture of engineering excellence, ownership, collaboration, and innovation.
Support hiring and talent development for critical infrastructure roles.
What We're Looking For 12+ years of experience in infrastructure engineering, distributed systems, cloud platforms, or AI infrastructure.
Proven experience leading large-scale infrastructure or platform engineering teams.
Deep understanding of GPU clusters, AI infrastructure, HPC, or large-scale distributed systems.
Strong expertise across
GPU Compute (NVIDIA ecosystem preferred)
High-performance networking (InfiniBand, RoCE, RDMA)
Kubernetes and container orchestration
Distributed storage systems
Infrastructure automation and observability
Linux systems and platform engineering
Experience designing highly available, scalable production infrastructure.
Robust architectural thinking with the ability to balance technical excellence and business priorities.
Excellent stakeholder management and cross-functional leadership skills.
Nice to Have Experience building AI factories, GPU cloud platforms, or large-scale inference infrastructure.
Familiarity with CUDA, NCCL, Slurm, Ray, Kubeflow, or similar AI infrastructure technologies.
Experience working with hyperscalers, cloud providers, or AI-first technology companies.
Exposure to multi-region or global infrastructure deployments.
Why Join
Nava? Build one of the world's leading AI infrastructure platforms.
Lead the engineering strategy behind next-generation GPU clusters and AI factories.
Work alongside world-class engineers solving some of the most challenging infrastructure problems in AI.
Shape the future of AI infrastructure from architecture through production at global scale.
Skills: leadership,automation,platforms,gpu,reliability,storage,architecture,cloud,infrastructure,cluster,drive
📌 Head of GPU Cluster Engineering (Bengaluru)
🏢 NAVA
📍 Bengaluru