11 Sep
|
NAVA
|
Bengaluru
About Nava
Nava is a neocloud company purpose-built for the AI era. We design, deploy, and operate large-scale GPU infrastructure and deliver inference-as-a-service to teams building the next generation of AI products. Our platform runs on NVIDIA GPU systems, high-performance RDMA fabrics, and a fully automated, software-defined operations model. Every engineer at Nava works close to the bare metal, on infrastructure built to keep GPUs available and models serving.
About The Team The Compute team owns Nava's GPU infrastructure end-to-end: from bare-metal provisioning of NVIDIA GPU systems through cluster software, RDMA integration, and production operation. We keep thousands of GPUs healthy and highly utilized so training and inference workloads run fast and reliably.
Responsibilities
- Provision and configure GPU nodes: OS, drivers, CUDA, GPU Operator, and RDMA connectivity.
- Write and maintain automation using Python, Ansible, and Terraform.
- Support cluster deployments, run health checks and benchmarks, and document results.
- Monitor GPU infrastructure, respond to alerts, and resolve or escalate issues.
- Maintain runbooks and operational documentation.
- Collaborate with cross-functional engineering teams to understand business and technical requirements and deliver software solutions.
Required Qualifications
- Strong hands-on experience with Kubernetes administration and architecture.
- Bare Metal GPU cluster deployment and management experience.
- Kubernetes networking, storage, and security.
- Infrastructure automation (Ansible, Terraform, etc.).
- Strong Linux systems expertise and working knowledge of the NVIDIA GPU stack (CUDA, drivers, GPU Operator) and RDMA (RoCEv2/InfiniBand).
- 2–4 years in systems/infrastructure engineering with robust Linux fundamentals and a learning-oriented mindset.
Preferred Qualifications
- NVIDIA GPU ecosystem exposure.
- AI/ML infrastructure experience.
- Monitoring tools such as Prometheus, Grafana, Dynatrace, Datadog, or Zabbix.
- Hybrid cloud and datacenter infrastructure experience.
- NCCL and NVLink tuning, Slurm, and distributed training/inference workloads.
Technology environment NVIDIA GPU systems (DGX/HGX-class, Blackwell), NVLink, CUDA, GPU Operator
- RoCEv2 / InfiniBand RDMA fabrics
- Kubernetes (networking, storage, security), Slurm
- Ansible, Terraform, Python; monitoring with Prometheus, Grafana, Dynatrace, Datadog, or Zabbix
- Linux at scale.
Skills: nvidia,kubernetes,cluster,rdma,metal,software,cuda,ansible,infrastructure,linux
📌 Network Engineer — GPU Infrastructure (Bengaluru)
🏢 NAVA
📍 Bengaluru