21 Sep
|
Xloud Technologies
|
New Delhi
21 Sep
Xloud Technologies
New Delhi
We’re looking for an engineer who can take ownership of designing, deploying and supporting GPU infrastructure for AI laboratories and research computing.
This role involves hands-on work across NVIDIA GPU servers, Lustre storage, high-speed networking and shared computing platforms, from initial architecture through deployment and ongoing support.
What you’ll work on
- Design and deploy Linux-based GPU and HPC clusters.
- Configure NVIDIA drivers, CUDA, MIG and GPU workload scheduling.
- Deploy and support Lustre, including high availability, performance tuning and recovery.
- Manage containerised AI workloads using Kubernetes and HPC scheduling with Slurm.
- Configure and troubleshoot high-speed Ethernet, RDMA/RoCE and storage connectivity.
- Automate provisioning and maintenance using Ansible, Python and Bash.
- Run benchmarks, diagnose performance bottlenecks and validate failover.
- Lead technical discussions with customers and OEMs, document deployments and support laboratory administrators.
What we’re looking for
- Proven experience deploying and supporting production HPC or GPU environments.
- Solid Linux administration and troubleshooting skills.
- Hands-on experience with NVIDIA GPUs and parallel storage, particularly Lustre.
- Understanding of compute, storage and networking dependencies.
- Ability to own technical delivery and resolve complex infrastructure incidents.
Additional experience with InfiniBand, BeeGFS, OpenPBS, NVIDIA AI Enterprise, Prometheus/Grafana, VMware or Proxmox would be valuable.
You’ll help build the infrastructure that researchers, faculty and students use for AI development and experimentation.
📌 Lead HPC & AI Infrastructure Engineer (New Delhi)
🏢 Xloud Technologies
📍 New Delhi