09 Aug
|
Larsen u0026 Toubro
|
Mumbai
09 Aug
Larsen u0026 Toubro
Mumbai
Job Purpose Provide high?throughput, consistent storage tiers (Scratch/HPS + Object) for large?scale training data ingest, checkpoints, and inference artifacts.
Roles & Responsibilities
- Implementation
Design/expand Lustre/BeeGFS HPS; NVMe?oF and Object (S3) tiers; align with AI dataflow and GDS. Establish namespace, OST/MDT layout, stripe/RAID policies; tiering for warm/cold datasets.
- Operations
Capacity/performance planning; rebalance and failover testing; automate snapshots and checkpoint retention. Proactive detection of hot spots and metadata contention; schema for small?file handling.
- Performance & Optimization
Tune RDMA paths, page cache, IO schedulers; validate end?to?end I/O profiles for LLM training/inference.
- Reliability & Incident
Lead P0/P1 critical incident response for enterprise-scale AI/ML storage infrastructure supporting NVIDIA GPU clusters, ensuring rapid service restoration and minimal impact to business-critical workloads. Act as the storage SME during major incidents involving BeeGFS, Lustre, GPFS (IBM Spectrum Scale), NFS, NVMe-oF, Parallel File Systems, and Object Storage platforms.
Perform deep-dive troubleshooting and resolution of storage performance degradation, metadata bottlenecks, I/O latency spikes, filesystem corruption, capacity exhaustion, and hardware failures.
Conduct comprehensive Root Cause Analysis (RCA) for storage-related outages affecting GPU training, inferencing, AI pipelines, and high-performance computing (HPC) workloads.
Backup/DR for critical datasets; test restore time and RPO/RTO regularly; corruption and split?brain handling.
- Security & Compliance
Multi?tenant encryption at rest and in transit, POSIX/ACLs, S3 policies, retention/legal holds. Experience & Educational Requirement BE/B-Tech or equivalent with Computer Science or Electronics & Communication
Certification- SNIA; vendor (NetApp/Dell/VAST) preferred.
Relevant Experience
- 7–12 years distributed storage; hands?on with Lustre/BeeGFS/Ceph, NVMe?oF, and S3 in GPU environments.
Tools / Tech
- Lustre/BeeGFS tooling, Ceph; GDS; Perf tools (fio, fs?digests); Prometheus/Grafana; Ansible.
📌 NVIDIA Storage Admin (Mumbai)
🏢 Larsen u0026 Toubro
📍 Mumbai