11 Sep
|
NAVA
|
Bengaluru
About Nava
Nava is a neocloud company purpose-built for the AI era. We design, deploy, and operate large-scale GPU infrastructure and deliver inference-as-a-service to teams building the next generation of AI products. Our platform runs on NVIDIA GPU systems, high-performance RDMA fabrics, and a fully automated, software-defined operations model. Every engineer at Nava works close to the metal, on infrastructure built to keep GPUs saturated and models serving.
About The Team The Compute team owns Nava's GPU infrastructure end-to-end: from bare-metal provisioning of NVIDIA GPU systems through cluster software, RDMA integration, and production operation. We keep thousands of GPUs healthy and highly utilized so training and inference workloads run fast and reliably.
Responsibilities
- Deploy and operate NVIDIA GPU clusters: bare-metal provisioning, node configuration, driver/CUDA/RDMA setup, and validation.
- Build automation in Python, Ansible, and Terraform to provision and maintain clusters and reduce manual toil.
- Extend Kubernetes using custom controllers and operators—design and develop custom Kubernetes controllers and operators.
- Build and maintain GitOps-based deployments, CI/CD pipelines, deployment automation, and infrastructure tooling.
- Integrate GPU nodes with RoCEv2/InfiniBand fabrics and validate end-to-end performance (NCCL, benchmarks).
- Lead new cluster turn-ups and acceptance testing; troubleshoot GPU, node, and interconnect issues.
- Improve monitoring and automated recovery to keep utilization high and MTTR low.
- Collaborate with Network and Storage teams and mentor junior engineers.
Required Qualifications
- Strong hands-on experience with Kubernetes administration and architecture.
- GPU cluster deployment and management experience.
- Deep understanding of Kubernetes networking, storage, and security.
- Experience with infrastructure automation tools (e.g., Ansible, Terraform).
- Strong Linux systems expertise and working knowledge of the NVIDIA GPU stack (CUDA, drivers, GPU Operator) and RDMA (RoCEv2/InfiniBand).
- 5–8 years in compute/systems infrastructure with hands-on GPU cluster experience; comfortable in an on-call production setting.
Preferred Qualifications
- Exposure to the NVIDIA GPU ecosystem.
- Experience with AI/ML infrastructure.
- Familiarity with monitoring tools such as Prometheus, Grafana, Dynatrace, Datadog, or Zabbix.
- Experience with hybrid cloud and datacenter infrastructure.
- Hands-on experience with NCCL and NVLink tuning, Slurm, and distributed training/inference workloads.
Technology environment NVIDIA GPU systems (DGX/HGX-class, Blackwell), NVLink, CUDA, GPU Operator
- RoCEv2 / InfiniBand RDMA fabrics
- Kubernetes (networking, storage, security), Slurm
- Ansible, Terraform, Python; monitoring with Prometheus, Grafana, Dynatrace, Datadog, or Zabbix
- Linux at scale.
Skills: metal,kubernetes,automation,nvidia,infiniband,cuda,cluster,infrastructure,rdma,ansible
📌 Senior Network Engineer — GPU Infrastructure (Bengaluru)
🏢 NAVA
📍 Bengaluru