18 Sep
|
DAMAC Digital
|
Bengaluru
18 Sep
DAMAC Digital
Bengaluru
AI Compute Engineer - Bengaluru, India
Someone has to design and tune the GPU clusters that actually run DAMAC AI's training and inference workloads. That's this role.
We're building one of the region's most ambitious AI compute footprints — NVIDIA B200, B300 and GB300 NVL72 clusters powering sovereign cloud and hyperscale AI services across the Middle East and Asia. You'll own architecture, deployment and performance across the full stack: bare metal, firmware, CUDA, NCCL, orchestration and multi-tenant scheduling.
What you'll do
- Architect GPU compute solutions for large-scale training and inference clusters
- Deploy and tune NVIDIA Base Command Manager, CUDA, NCCL, cuDNN, TensorRT, DCGM, MIG
- Integrate with InfiniBand/RoCE fabrics and onboard Slurm, Kubernetes, Run:ai
- Benchmark with NCCL tests, MLPerf and HPL,
and chase down every bottleneck
- Own firmware/driver upgrade cycles and lead vendor engagements with NVIDIA, Supermicro, Dell, HPE
What you bring
- 7+ years in HPC, AI or GPU infrastructure, 3+ years hands-on with NVIDIA GPU clusters
- Deep NVIDIA stack expertise: CUDA, NCCL, cuDNN, TensorRT, DCGM
- InfiniBand/RoCE networking knowledge and Slurm/Kubernetes/Run:ai experience
- Advanced Linux, Python/Bash, Ansible or Terraform
- Bonus: NVIDIA certification, LLM training at scale, liquid-cooled GB300 NVL72 experience
If you'd rather be tuning NCCL collectives than sitting in another status meeting — this is your seat.
📌 AI Compute Engineer (Bengaluru)
🏢 DAMAC Digital
📍 Bengaluru