12 Sep
|
NAVA
|
Bengaluru
About Nava
Nava is a neocloud company purpose-built for the AI era. We design, deploy, and operate large-scale GPU infrastructure and deliver inference-as-a-service to teams building the next generation of AI products. Our platform runs on NVIDIA GPU systems, high-performance RDMA fabrics, and a fully automated, software-defined operations model. Every engineer at Nava works close to the metal, on infrastructure built to keep GPUs saturated and models serving.
About the team The Network Engineering team owns the full lifecycle of Nava's AI datacenter fabric, from architecture and design through deployment, automation, and production operation. We build the lossless, high-throughput networks that carry GPU-to-GPU traffic for distributed training and low-latency inference. This is a hands-on team where the network is treated as code—and where your expertise will shape the foundation of AI infrastructure at scale.
Responsibilities
· Own the architecture and high-level design of Nava's GPU cluster fabrics, including RoCEv2 backend networks, frontend networks, and datacenter interconnect.
· Define fleet-wide standards for BGP and EVPN-VXLAN leaf-spine topologies, congestion control (PFC/ECN tuning), and lossless RDMA transport.
· Set the technical direction for network automation, decomposing high-level architecture into detailed, buildable designs.
· Lead cross-functional initiatives with Compute, Storage, SRE, and security teams to deliver fabrics that never bottleneck training or inference.
· Drive root-cause analysis on the hardest fabric performance and reliability problems and establish preventive practices.
· Mentor Senior and Network Engineers, review designs, and raise the technical bar across the team.
· Remain hands-on: lead cluster turn-ups, acceptance testing, RoCEv2 implementation, incident response,
and PFC/ECN tuning under real load, alongside architecture and strategy.
Required qualifications
· 8–10 years designing and operating datacenter networks at scale, with protocol-level architecture authority and a track record leading complex, multi-vendor designs.
· Strong hands-on command of core networking protocols: BGP, OSPF, IS-IS, TCP/IP, IPv4 and IPv6, DNS, DHCP, and MPLS.
· Experience with networking protocols such as TCP/IP, VPN, DNS, DHCP, and SSL/TLS.
· Solid understanding of datacenter fabric concepts (leaf-spine, EVPN-VXLAN) and RDMA fabrics (RoCEv2 and InfiniBand), including lossless Ethernet.
· Automation skills: Python plus Ansible and/or Terraform.
· Hands-on experience with RDMA fabrics (RoCEv2 and/or InfiniBand), including congestion management for lossless Ethernet and IB fabrics.
· Comfortable operating in a fast-paced, on-call production setting.
Preferred qualifications
· Experience with NVIDIA networking (Spectrum-X, Quantum InfiniBand, BlueField DPUs) and NCCL traffic patterns.
· Experience operating GPU clusters for large-scale distributed training or inference.
· Familiarity with network telemetry, streaming analytics, and closed-loop automation.
· Relevant certifications (e.g., CCNP or vendor equivalents).
· Experience with Cisco and/or Arista platforms is a strong plus.
Technology environment
NVIDIA GPU systems (DGX/HGX-class, Blackwell), NVLink, NVIDIA networking (Spectrum/Quantum, BlueField DPUs); RoCEv2 and InfiniBand RDMA lossless fabrics; BGP, OSPF, IS-IS, EVPN-VXLAN, MPLS; TCP/IP, IPv4/IPv6, DNS, DHCP, VPN, SSL/TLS; Ansible, Terraform, Python; Linux at scale.
Ready to help build the networking backbone for the next wave of AI innovation? We’d love to hear from you. Apply today—and let’s shape the future of AI infrastructure, together.
📌 Principal Network Engineer — AI Datacenter (Bengaluru)
🏢 NAVA
📍 Bengaluru