11 Sep
|
NAVA
|
Bengaluru
About Nava
Nava is a neocloud company purpose-built for the AI era. We design, deploy, and operate large-scale GPU infrastructure and deliver inference-as-a-service to teams building the next generation of AI products. Our platform runs on NVIDIA GPU systems, high-performance RDMA fabrics, and a fully automated, software-defined operations model. Every engineer at Nava works close to the metal, on infrastructure built to keep GPUs saturated and models serving.
About The Team The Network Engineering team owns the full lifecycle of Nava's AI datacenter fabric, from architecture and design through deployment, automation, and production operation. We build the lossless, high-throughput networks that carry GPU-to-GPU traffic for distributed training and low-latency inference. This is a hands-on team where the network is treated as code.
Responsibilities
- Deploy, configure, and support leaf-spine datacenter networks running BGP and EVPN-VXLAN.
- Help implement and troubleshoot RoCEv2 / InfiniBand RDMA connectivity for GPU clusters.
- Write and maintain automation using Python, Ansible, and Terraform to reduce manual work.
- Support cluster turn-ups, run validation and network benchmarks, and document results.
- Monitor the fabric, respond to alerts, and resolve or escalate incidents.
- Maintain accurate network documentation and runbooks.
Required Qualifications
- 2–4 years in network engineering,
with a solid grounding in routing/switching.
- Strong hands-on command of core networking protocols: BGP, OSPF, IS-IS, TCP/IP (IPv4/IPv6), DNS, DHCP, MPLS, and SSL/TLS.
- Experience with datacenter fabric architectures (leaf-spine, EVPN-VXLAN) and RDMA technologies (RoCEv2 and InfiniBand), including lossless Ethernet configuration and troubleshooting.
- Automation proficiency: Python, Ansible, and/or Terraform.
- Solid troubleshooting instincts and a proactive, learning-oriented mindset.
Preferred Qualifications
- Experience with NVIDIA networking (Spectrum-X, Quantum InfiniBand, BlueField DPUs) and NCCL traffic patterns.
- Experience operating GPU clusters for large-scale distributed training or inference.
- Familiarity with network telemetry, streaming analytics, and closed-loop automation.
- Relevant certifications (e.g., CCNP or vendor equivalents).
- Experience with Cisco and/or Arista platforms is a strong plus.
Technology environment
NVIDIA GPU systems (DGX/HGX-class, Blackwell), NVLink, NVIDIA networking (Spectrum/Quantum, BlueField DPUs)
- RoCEv2 and InfiniBand RDMA lossless fabrics
- BGP, OSPF, IS-IS, EVPN-VXLAN, MPLS
- TCP/IP, IPv4/IPv6, DNS, DHCP, VPN, SSL/TLS
- Ansible, Terraform, Python
- Linux at scale.
Skills: automation,nvidia,infiniband,ip,bgp,tcp/ip,python,rdma,ansible,networking
📌 Network Engineer — AI Datacenter (Bengaluru)
🏢 NAVA
📍 Bengaluru