19 Sep
|
DAMAC Digital
|
Bengaluru
19 Sep
DAMAC Digital
Bengaluru
AI Compute Engineer - Bengaluru, India
Someone has to design and tune the GPU clusters that actually run DAMAC AI's training and inference workloads. That's this role.
We're building one of the region's most ambitious AI compute footprints — NVIDIA B200, B300 and GB300 NVL72 clusters powering sovereign cloud and hyperscale AI services across the Middle East and Asia. You'll own architecture, deployment and performance across the full stack: bare metal, firmware, CUDA, NCCL, orchestration and multi-tenant scheduling.
What you'll do
Architect GPU compute solutions for large-scale training and inference clusters
Deploy and tune NVIDIA Base Command Manager, CUDA, NCCL, cuDNN, TensorRT, DCGM, MIG
Integrate with InfiniBand/RoCE fabrics and onboard Slurm, Kubernetes, Run:ai
Benchmark with NCCL tests, MLPerf and HPL,
and chase down every bottleneck
Own firmware/driver upgrade cycles and lead vendor engagements with NVIDIA, Supermicro, Dell, HPE
What you bring
7+ years in HPC, AI or GPU infrastructure, 3+ years hands-on with NVIDIA GPU clusters
Deep NVIDIA stack expertise: CUDA, NCCL, cuDNN, TensorRT, DCGM
InfiniBand/RoCE networking knowledge and Slurm/Kubernetes/Run:ai experience
Advanced Linux, Python/Bash, Ansible or Terraform
Bonus: NVIDIA certification, LLM training at scale, liquid-cooled GB300 NVL72 experience
If you'd rather be tuning NCCL collectives than sitting in another status meeting — this is your seat.
📌 Ai Compute Engineer Bengaluru
🏢 DAMAC Digital
📍 Bengaluru