Principal Network Engineer — GPU Infrastructure (Bengaluru)

Principal Network Engineer — GPU Infrastructure (Bengaluru)

11 Sep
|
NAVA
|
Bengaluru

11 Sep

NAVA

Bengaluru

About Nava

Nava is a neocloud company built for the AI era. We design, deploy, and operate large-scale GPU infrastructure—and deliver inference-as-a-service to teams building next-generation AI products. Our platform runs on NVIDIA GPUs, high-performance networking (like RoCEv2 or InfiniBand), and a fully automated, software-defined operations model. Engineers at Nava work closely with the hardware to keep GPUs fully utilized and models serving efficiently.

About The Team The Compute team owns Nava’s GPU infrastructure end-to-end—from installing and configuring NVIDIA GPU systems, to managing cluster software, networking, and day-to-day operations.

Our goal: keep thousands of GPUs healthy, responsive, and running at peak performance for training and inference workloads.

Responsibilities

- Design and evolve the architecture for large-scale NVIDIA GPU clusters connected via NVLink and high-speed networking (RoCEv2/InfiniBand).
- Define standards for bare-metal provisioning, GPU node setup (drivers, CUDA, GPU Operator, RDMA), and cluster lifecycle automation.
- Lead automation efforts across the fleet using Python, Ansible, and Terraform to deploy, configure, and maintain GPU clusters at scale.
- Build and maintain practices for GPU health monitoring, fault detection, and rapid recovery when nodes fail.




- Collaborate with Network and Storage teams to ensure data paths are optimized and GPUs stay fully utilized.
- Mentor both senior and junior engineers—and help set technical standards for the team.

Qualifications

- Hands-on experience with Kubernetes administration and architecture.
- Experience deploying and managing GPU clusters in production environments.
- Strong understanding of Kubernetes networking, storage, and security.
- Experience building infrastructure automation using tools like Ansible and Terraform.
- Deep Linux systems knowledge—and familiarity with the NVIDIA GPU stack (CUDA, drivers, GPU Operator) and RDMA (RoCEv2/InfiniBand).
- 8–10 years in infrastructure engineering—with recent experience in compute infrastructure and a history of leading architecture for large-scale systems.
- Experience with NVIDIA GPU systems (e.g., DGX/HGX-class, Blackwell), NVLink, CUDA, GPU Operator, and RDMA fabrics.
- Experience with monitoring tools such as Prometheus, Grafana, Dynatrace, Datadog, or Zabbix.
- Familiarity with Slurm, NCCL, NVLink tuning, and distributed training/inference workloads is a plus.

Preferred Qualifications

- Experience in AI/ML infrastructure.
- Experience with hybrid cloud and datacenter infrastructure.

Skills: nvidia,kubernetes,cluster,rdma,infiniband,cuda,architecture,ansible,infrastructure,automation

📌 Principal Network Engineer — GPU Infrastructure (Bengaluru)
🏢 NAVA
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: principal network engineer — gpu infrastructure (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: principal network engineer — gpu infrastructure (bengaluru) / bengaluru