We’re looking for an engineer who can take ownership of designing, deploying and supporting GPU infrastructure for AI laboratories and research computing.
This role involves hands-on work across NVIDIA GPU servers, Lustre storage, high-speed networking and shared computing platforms, from initial architecture through deployment and ongoing support.
What you’ll work on
- Design and deploy Linux-based GPU and HPC clusters.
- Configure NVIDIA drivers, CUDA, MIG and GPU workload scheduling.
- Deploy and support Lustre, including high availability, performance tuning and recovery.
- Manage containerised AI workloads using Kubernetes and HPC scheduling with Slurm.
- Configure and troubleshoot high-speed Ethernet, RDMA/RoCE and storage connectivity.
- Automate provisioning and maintenance using Ansible, Python and Bash.
- Run benchmarks, diagnose performance bottlenecks and validate failover.
- Lead technical discussions with customers and OEMs, document deployments and support laboratory administrators.
What we’re looking for
- Proven experience deploying and supporting production HPC or GPU environments.
- Robust Linux administration and troubleshooting skills.
- Hands-on experience with NVIDIA GPUs and parallel storage, particularly Lustre.
- Understanding of compute, storage and networking dependencies.
- Ability to own technical delivery and resolve complex infrastructure incidents.
Additional experience with InfiniBand, BeeGFS, OpenPBS, NVIDIA AI Enterprise, Prometheus/Grafana, VMware or Proxmox would be valuable.
You’ll help build the infrastructure that researchers, faculty and students use for AI development and experimentation.
📌 Lead HPC & AI Infrastructure Engineer (New Delhi)
🏢 Xloud Technologies
📍 New Delhi
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.