26 Aug
|
Yotta Infrastructure
|
Mumbai
26 Aug
Yotta Infrastructure
Mumbai
Job Summary
The role is responsible for designing, deploying, and operating large-scale AI infrastructure, including GPU compute, high-performance networking, storage, and orchestration platforms. The incumbent will lead the architecture of secure, scalable, and high-availability AI clusters to support distributed training, fine-tuning, and inference workloads, while ensuring optimal performance, reliability, and compliance.
Experience
- 10+ years in systems engineering, network engineering, cloud infrastructure, or datacenter design
Key Responsibilities
- AI Systems Architecture (Compute, GPU, OS): Design and deploy large-scale GPU clusters (H100, H200, GB200 and GB300) for distributed training and inference.
- Architect multi-node GPU systems using: NVLink/NVSwitch; PCIe Gen5
- Define OS, kernel, driver, and runtime configurations optimized for AI workloads (CUDA/ROCm, NCCL, UCX, OFED).
- Develop high-performance compute blueprints for diverse use cases: training, fine-tuning, retrieval, and batch inference.
- High-Performance Networking: Architect AI fabric networks including InfiniBand HDR/NDR/XDR/SPX; RoCEv2 / RDMA; 100/200/400/800 Gbps Ethernet fabrics.
- Design low-latency, high-bandwidth topologies (fat-tree, dragonfly+, multi-plane architectures).
- Plan and tune inter-node communication for distributed AI training (NCCL, MPI, UCX).
- Implement network segmentation, isolation, and multi-tenant security for AI compute clusters.
- Storage Data Pipeline Infrastructure: Architect high-throughput storage solutions for AI: Parallel file systems (Lustre, BeeGFS, IBM Spectrum Scale); Cloud-native high-performance storage (FSx for Lustre,
Azure ANF, GCS Filestore High Scale); NVMe, NVMe-over-Fabrics, object storage.
- Optimize data pipelines for large-scale dataset ingestion, feature extraction, checkpointing, and streaming.
- Platform Integration Orchestration: Integrate systems with Kubernetes GPU environments (EKS/AKS/GKE, K8s on-prem, Kueue, Volcano).
- Design infrastructure to support distributed training frameworks: PyTorch DDP; DeepSpeed; Ray Train; JAX / TPU alternatives.
- Enable robust scheduling, multi-tenancy, and job orchestration.
- Reliability, Monitoring Performance Optimization: Implement monitoring for GPU utilization, network telemetry, I/O performance, and cluster health (Prometheus, Grafana, DCGM, NetQ).
- Conduct performance tuning across: NIC/driver stack; GPU topology; Storage throughput; Network congestion management (ECN, PFC, QoS).
- Design systems for high availability, resilience, and disaster recovery.
- Security Compliance (Infra-Level): Implement hardware-level and network-level security controls IAM, RBAC, ACLs, segmentation, encryption in transit.
- Architect secure multi-tenant GPU environments, including confidential computing where supported.
- Ensure system compliance with SOC2, ISO 27001, or industry-specific security frameworks.
Key Skills
- GPU clusters
- NVLink/NVSwitch
- PCIe Gen5
- CUDA
- ROCm
- NCCL
- Kubernetes
- PyTorch DDP
- DeepSpeed
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Lead Solution Architect (Mumbai)
🏢 Yotta Infrastructure
📍 Mumbai