14 Aug
|
NAVA
|
Bengaluru
About Nava Nava is building next-generation AI infrastructure and inference platforms at global scale. We're looking for a Principal Engineer - Cluster Deployment to lead the successful deployment of large-scale GPU clusters across global data centres. This is a high-impact execution role responsible for taking GPU clusters from hardware delivery to fully validated, production-ready infrastructure.
You'll own deployment execution across sites, vendors, and cross-functional teams, ensuring every cluster is delivered on time, meets quality standards, and is ready for customer workloads.
What You'll Do End-to-End Cluster Deployment
Own the deployment of GPU clusters from rack arrival through to tenant-ready production capacity.
Plan, coordinate, and execute deployment activities across multiple global data centre locations.
Drive installation, rack integration, power-up, hardware bring-up, firmware validation, and production readiness.
Ensure deployments are completed safely, efficiently, and within agreed timelines.
Deployment Execution
Lead cluster bring-up activities, including hardware validation, network configuration, storage integration, and platform initialization.
Oversee structured cabling, power validation, firmware upgrades, BIOS configuration, and infrastructure readiness.
Coordinate deployment activities across internal engineering teams, hardware vendors, systems integrators, and data centre partners.
Resolve deployment blockers and drive rapid issue resolution during implementation.
Validation & Quality Assurance
Develop and execute deployment validation procedures and acceptance criteria.
Lead burn-in testing, stress testing,
hardware diagnostics, and production readiness validation.
Ensure clusters meet defined performance, stability, and reliability benchmarks before customer handover.
Own first-pass acceptance and minimize deployment rework through robust quality processes.
Vendor & Site Management
Act as the primary technical lead during in office deployments.
Manage relationships with OEMs, contract manufacturers, data centre operators, and installation partners.
Ensure deployment standards are consistently followed across all locations.
Drive continuous improvement across deployment processes, documentation, and execution methodologies.
Cross-Functional Collaboration
Partner with GPU Cluster Engineering, Platform Engineering, Networking, SRE, Supply Chain, and Data Centre Operations teams.
Ensure seamless transition from deployment to production operations.
Support troubleshooting of deployment issues and coordinate engineering fixes where required.
Operational Excellence
Develop deployment playbooks, SOPs, checklists, and automation to improve deployment consistency.
Track deployment metrics and identify opportunities to reduce deployment timelines and improve quality.
Drive lessons learned reviews following every deployment.
Success Metrics You Will Be Measured On Time-to-live for new GPU clusters
First-pass deployment acceptance rate
Deployment quality and reliability
Reduction in deployment rework
Deployment schedule adherence
Production readiness at handover
Deployment process standardization and automation Qualifications Required Qualifications 10+ years of experience in infrastructure deployment, data centre engineering, HPC, cloud infrastructure, or systems engineering.
Proven experience deploying large-scale compute infrastructure, GPU clusters, HPC systems, or cloud platforms.
Deep technical expertise in
Server and rack deployment
GPU hardware platforms
High-speed networking (InfiniBand, Ethernet, RDMA)
Structured cabling and power systems
Linux systems administration
Firmware, BIOS, and hardware lifecycle management
Infrastructure validation and production readiness
Experience coordinating cross-functional deployment projects across multiple sites.
Strong troubleshooting and problem-solving skills in complex infrastructure environments.
Willingness to travel to domestic and international deployment locations.
Preferred Qualifications Experience with NVIDIA DGX, HGX, or similar GPU platforms.
Familiarity with Kubernetes, cluster provisioning, automation, and Infrastructure-as-Code.
Experience working with hyperscalers, AI infrastructure providers, or data centre operators.
Exposure to HPC, AI factories, or large-scale inference platforms.
Skills: drive,platforms,infrastructure,data,gpu,production readiness,automation,cluster,validation,firmware,readiness
📌 Principal Engineer – Cluster Deployment (Bengaluru)
🏢 NAVA
📍 Bengaluru