Job Summary
Senior Specialist Consultant Module Lead P5 supporting AI Datacenter under the Joint GTM capability enhancement track. Designs, deploys, supports GPU clusters and AI infra platforms, optimizes AIML workload performance. Operates at L3 level within the System Core Engg SME L2 L3 L4 tower with 9 to 12 years of relevant experience. Embedded full-time within delivery teams to build, run and scale the sovereign compute AI stack ensuring SLA adherence, documentation and knowledge transfer. Collaborates closely with stakeholders; follows change/incident management processes and contributes to continuous improvement, automation and regulatory compliance readiness across the engagement. GPU Cluster Design; AI Infra Deployment; AIML Workload Optimization; NVIDIA Platform Expertise; CUDA/NIM Stack; DGX Reference Architectures; MIG/GPU Partitioning; Performance Benchmarking.
Responsibilities
- Supporting AI Datacenter under the Joint GTM capability enhancement track.
- Designs, deploys, supports GPU clusters and AI infra platforms.
- Optimizes AIML workload performance.
- Operates at L3 level within the System Core Engg SME L2 L3 L4 tower with 9 to 12 years of relevant experience.
- Embedded full time within delivery teams to build, run and scale the sovereign compute/AI stack ensuring SLA adherence, documentation and knowledge transfer.
- Collaborates closely with stakeholders; follows change/incident management processes and contributes to continuous improvement, automation and regulatory compliance readiness across the engagement.
- GPU Cluster Design.
- AI Infra Deployment.
- AIML Workload Optimization.
- NVIDIA Platform Expertise.
- CUDA/NIM Stack.
- DGX Reference Architectures.
- MIG/GPU Partitioning.
- Performance Benchmarking.
Requirements
- 9 to 12 years of relevant experience.
- Embedded full-time within delivery teams to build, run and scale the sovereign compute/AI stack ensuring SLA adherence.
- Operates at L3 level within the System Core Engg SME L2 L3 L4 tower.
- Collaborates closely with stakeholders; follows change/incident management processes.
- Documentation and knowledge transfer.
- Regulatory compliance readiness across the engagement.
Skills
- GPU Cluster Design
- AI Infra Deployment
- AIML Workload Optimization
- NVIDIA Platform Expertise
- CUDA/NIM Stack
- DGX Reference Architectures
- MIG/GPU Partitioning
- Performance Benchmarking
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Associate Principal - System Management (Chennai)
🏢 LTM
📍 Chennai