13 Aug
|
NAVA
|
Bengaluru
About Nava
Nava is building next-generation AI infrastructure and inference platforms at global scale. We're looking for a Principal Engineer - Cluster Deployment to lead the successful deployment of large-scale GPU clusters across global data centres.
This is a high-impact execution role responsible for taking GPU clusters from hardware delivery to fully validated, production-ready infrastructure. You'll own deployment execution across sites, vendors, and cross-functional teams, ensuring every cluster is delivered on time, meets quality standards, and is ready for customer workloads.
What You'll Do
- End-to-End Cluster Deployment
- Own the deployment of GPU clusters from rack arrival through to tenant-ready production capacity.
- Plan, coordinate, and execute deployment activities across multiple global data centre locations.
- Drive installation, rack integration, power-up, hardware bring-up, firmware validation, and production readiness.
- Ensure deployments are completed safely, efficiently, and within agreed timelines.
- Deployment Execution
- Lead cluster bring-up activities, including hardware validation, network configuration, storage integration, and platform initialization.
- Oversee structured cabling, power validation, firmware upgrades, BIOS configuration, and infrastructure readiness.
- Coordinate deployment activities across internal engineering teams, hardware vendors, systems integrators, and data centre partners.
- Resolve deployment blockers and drive rapid issue resolution during implementation.
- Validation & Quality Assurance
- Develop and execute deployment validation procedures and acceptance criteria.
- Lead burn-in testing, stress testing,
hardware diagnostics, and production readiness validation.
- Ensure clusters meet defined performance, stability, and reliability benchmarks before customer handover.
- Own first-pass acceptance and minimize deployment rework through robust quality processes.
- Vendor & Site Management
- Act as the primary technical lead during on-site deployments.
- Manage relationships with OEMs, contract manufacturers, data centre operators, and installation partners.
- Ensure deployment standards are consistently followed across all locations.
- Drive continuous improvement across deployment processes, documentation, and execution methodologies.
- Cross-Functional Collaboration
- Partner with GPU Cluster Engineering, Platform Engineering, Networking, SRE, Supply Chain, and Data Centre Operations teams.
- Ensure seamless transition from deployment to production operations.
- Support troubleshooting of deployment issues and coordinate engineering fixes where required.
- Operational Excellence
- Develop deployment playbooks, SOPs, checklists, and automation to improve deployment consistency.
- Track deployment metrics and identify opportunities to reduce deployment timelines and improve quality.
- Drive lessons learned reviews following every deployment.
Success Metrics
You Will Be Measured On
- Time-to-live for current GPU clusters
- First-pass deployment acceptance rate
- Deployment quality and reliability
- Reduction in deployment rework
- Deployment schedule adherence
- Production readiness at handover
- Deployment process standardization and automation
Qualifications
Required Qualifications
- 10+ years of experience in infrastructure deployment, data centre engineering, HPC, cloud infrastructure, or systems engineering.
- Proven experience deploying large-scale compute infrastructure, GPU clusters, HPC systems, or cloud platforms.
- Deep technical expertise in:
- Server and rack deployment
- GPU hardware platforms
- High-speed networking (InfiniBand, Ethernet, RDMA)
- Structured cabling and power systems
- Linux systems administration
- Firmware, BIOS, and hardware lifecycle management
- Infrastructure validation and production readiness
- Experience coordinating cross-functional deployment projects across multiple sites.
- Strong troubleshooting and problem-solving skills in complex infrastructure environments.
- Willingness to travel to domestic and international deployment locations.
Preferred Qualifications
- Experience with NVIDIA DGX, HGX, or similar GPU platforms.
- Familiarity with Kubernetes, cluster provisioning, automation, and Infrastructure-as-Code.
- Experience working with hyperscalers, AI infrastructure providers, or data centre operators.
- Exposure to HPC, AI factories, or large-scale inference platforms.
Skills: drive,platforms,infrastructure,data,gpu,production readiness,automation,cluster,validation,firmware,readiness
📌 Principal Engineer Cluster Deployment (Bengaluru)
🏢 NAVA
📍 Bengaluru