Network Engineer — GPU Infrastructure (Bengaluru)

Network Engineer — GPU Infrastructure (Bengaluru)

12 Sep
|
NAVA
|
Bengaluru

12 Sep

NAVA

Bengaluru

About Nava Nava is a neocloud company purpose-built for the AI era. We design, deploy, and operate large-scale GPU infrastructure and deliver inference-as-a-service to teams building the next generation of AI products. Our platform runs on NVIDIA GPU systems, high-performance RDMA fabrics, and a fully automated, software-defined operations model. Every engineer at Nava works close to the bare metal, on infrastructure built to keep GPUs available and models serving.

About The Team The Compute team owns Nava's GPU infrastructure end-to-end: from bare-metal provisioning of NVIDIA GPU systems through cluster software, RDMA integration, and production operation. We keep thousands of GPUs healthy and highly utilized so training and inference workloads run fast and reliably.

Responsibilities

- Provision and configure GPU nodes: OS, drivers, CUDA, GPU Operator, and RDMA connectivity.
- Write and maintain automation using Python, Ansible, and Terraform.
- Support cluster deployments, run health checks and benchmarks, and document results.
- Monitor GPU infrastructure, respond to alerts, and resolve or escalate issues.
- Maintain runbooks and operational documentation.
- Collaborate with cross-functional engineering teams to understand business and technical requirements and deliver software solutions.





Required Qualifications

- Solid hands-on experience with Kubernetes administration and architecture.
- Bare Metal GPU cluster deployment and management experience.
- Kubernetes networking, storage, and security.
- Infrastructure automation (Ansible, Terraform, etc.).
- Strong Linux systems expertise and working knowledge of the NVIDIA GPU stack (CUDA, drivers, GPU Operator) and RDMA (RoCEv2/InfiniBand).
- 2–4 years in systems/infrastructure engineering with strong Linux fundamentals and a learning-oriented mindset.

Preferred Qualifications

- NVIDIA GPU ecosystem exposure.
- AI/ML infrastructure experience.
- Monitoring tools such as Prometheus, Grafana, Dynatrace, Datadog, or Zabbix.
- Hybrid cloud and datacenter infrastructure experience.
- NCCL and NVLink tuning, Slurm, and distributed training/inference workloads.

Technology environment NVIDIA GPU systems (DGX/HGX-class, Blackwell), NVLink, CUDA, GPU Operator; RoCEv2 / InfiniBand RDMA fabrics; Kubernetes (networking, storage, security), Slurm; Ansible, Terraform, Python; monitoring with Prometheus, Grafana, Dynatrace, Datadog, or Zabbix; Linux at scale. Skills: nvidia,kubernetes,cluster,rdma,metal,software,cuda,ansible,infrastructure,linux

📌 Network Engineer — GPU Infrastructure (Bengaluru)
🏢 NAVA
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: network engineer — gpu infrastructure (bengaluru) / bengaluru