06 Aug
|
Larsen and Toubro (L&T)
|
Chennai
06 Aug
Larsen and Toubro (L&T)
Chennai
Job Purpose
Provide highthroughput, consistent storage tiers (Scratch/HPS + Object) for largescale training data ingest, checkpoints, and inference artifacts.
Role Description
Key Responsibilities
- Implementation
- Design/expand Lustre/BeeGFS HPS; NVMeoF and Object (S3) tiers; align with AI dataflow and GDS.
- Establish namespace, OST/MDT layout, stripe/RAID policies; tiering for warm/cold datasets.
- Operations
- Capacity/performance planning; rebalance and failover testing; automate snapshots and checkpoint retention.
- Proactive detection of hot spots and metadata contention; schema for smallfile handling.
- Performance & Optimization
- Tune RDMA paths, page cache, IO schedulers; validate endtoend I/O profiles for LLM training/inference.
- Reliability & Incident
- Lead P0/P1 critical incident response for enterprise-scale AI/ML storage infrastructure supporting NVIDIA GPU clusters, ensuring rapid service restoration and minimal impact to business-critical workloads.
- Act as the storage SME during major incidents involving BeeGFS, Lustre, GPFS (IBM Spectrum Scale), NFS, NVMe-oF, Parallel File Systems, and Object Storage platforms.
- Perform deep-dive troubleshooting and resolution of storage performance degradation, metadata bottlenecks, I/O latency spikes, filesystem corruption, capacity exhaustion, and hardware failures.
- Conduct comprehensive Root Cause Analysis (RCA) for storage-related outages affecting GPU training, inferencing, AI pipelines, and high-performance computing (HPC) workloads.
- Backup/DR for critical datasets; test restore time and RPO/RTO regularly; corruption and splitbrain handling.
- Security & Compliance
- Multitenant encryption at rest and in transit, POSIX/ACLs, S3 policies, retention/legal holds.
Experience & Educational Requirements
Qualifications and Experience
EDUCATIONAL QUALIFICATIONS: (degree, training, or certification required)
BE/B-Tech or equivalent with Computer Science or Electronics & Communication
RELEVANT EXPERIENCE: (no. of years of technical, functional, and/or leadership experience or specific exposure required)
- 712 years distributed storage; handson with Lustre/BeeGFS/Ceph, NVMeoF, and S3 in GPU environments.
Tools / Tech
Lustre/BeeGFS tooling, Ceph; GDS; Perf tools (fio, fsdigests); Prometheus/Grafana; Ansible.
Certifications
SNIA; vendor (NetApp/Dell/VAST) preferred.
KPIs
Sustained/peak throughput, metadata ops latency, job I/O stall rate, restore drill success.
Work Mode
Hybrid; participates in maintenance windows and DR exercises.
📌 Storage Engineer (Chennai)
🏢 Larsen and Toubro (L&T)
📍 Chennai