06 Aug
|
Larsen and Toubro (L&T)
|
Chennai
06 Aug
Larsen and Toubro (L&T)
Chennai
Job Purpose
Build and operate largescale GPU compute pods to deliver predictable, highthroughput, lowlatency training and inference services across 10K GPU cluster.
Role Description
Key Responsibilities
- Implementation
- Stand up multipod GPU clusters (rack/power/cooling layouts; TOR/leaf connectivity; IB/Ethernet host configs).
- Implement GPU partitioning (MIG/vGPU profiles) and quota policies for multitenant environments.
- Integrate cluster schedulers (Slurm/Kubernetes) with GPU device plugins, node feature discovery, and accounting/quotas.
- Operations
- Own day2 operations across firmware/driver/DCGM/NVML lifecycles; execute change windows with zero/low downtime.
- Capacity planning (GPU/CPU/Memory/NIC) and binpacking strategies for heterogeneous GPU SKUs.
- Performance & Optimization
- Tune NCCL/UCX, GPU clocks/persistence, GPU Direct Storage, NUMA/locality, and CUDA runtime parameters.
- Drive benchmarking and acceptance (HPL, HPLAI, MLPerflike internal suites); track perf regressions with SLOs.
- Reliability & Incident
- Lead P0/P1 incident response for GPU, CUDA, or scheduler issues; perform rootcause and preventative actions.
- Act as the primary technical lead during production outages, coordinating cross-functional teams across Networking, GPU Operations, Platform Engineering, Storage, and Application teams to restore services within SLA targets.
- Perform detailed Root Cause Analysis (RCA) for network, GPU, CUDA, NCCL, RDMA, and scheduler-related failures, identifying underlying causes and implementing preventive and corrective actions.
- Collaborate with platform and GPU engineering teams to resolve issues impacting CUDA jobs, Kubernetes scheduling, Slurm workload management, GPU resource allocation, and large-scale AI training environments.
- Define golden images; implement node remediation (cordon/drain/reimage) and autohealing workflows.
- Security & Compliance
- Enforce GPU tenancy isolation (MIG, cgroup/device cgroup, mpsd), secure drivers/containers, SBOM and image scanning.
- Documentation & Enablement
- Publish runbooks, performance baselines, and application tuning guides for LLM training and inference.
Experience & Educational Requirements
Qualifications and Experience
EDUCATIONAL QUALIFICATIONS: (degree, training, or certification required)
BE/B-Tech or equivalent with Computer Science or Electronics & Communication
RELEVANT EXPERIENCE: (no. of years of technical, functional, and/or leadership experience or specific exposure required)
- 712 years building/operating HPC/AI GPU clusters at scale; deep CUDA/MIG/vGPU expertise.
- Proven with H100/Bseries class GPUs, Slurm and/or Kubernetes device scheduling; NCCL/UCX performance tuning.
Tools / Tech
1. CUDA, NCCL, cuDNN, TensorRT; 2. DCGM/NVML, nvidia-smi; 3. Slurm, K8s (device plugin, DaemonSets, NFD); 4. Helm/ArgoCD; 6. GPUDirect Storage; 7. Prometheus/Grafana; 8. ELK/Splunk.
Certifications
NVIDIA Certified Qualified (AI Infrastructure); Linux (RHCE/LPIC) preferred.
KPIs
GPU utilization %, queue waittime SLOs, failedjob rate, perf baseline adherence, MTTD/MTTR for GPU incidents.
Work Mode
Onsite/Hybrid; participates in change windows and critical afterhours events.
📌 Nvidia GPU Infrastructure Engineer / HPC Engineer (Chennai)
🏢 Larsen and Toubro (L&T)
📍 Chennai