Designs and architect end-to-end AI Cloud platforms with a focus on security, cost-efficiency, and performance. This position involves direct client engagement to translate requirements into technical Solution, encompassing GPU infrastructure rightsizing and optimal model selection. We are looking for a cloud expert with a demonstrated ability to transition complex AI models from concept to large-scale production. The ideal candidate brings extensive experience in AI/Cloud ecosystems and a successful track record of architecting and managing production-grade, large-scale AI platforms.
Role Summary
Key Responsibilities
- Translate business requirements into scalable, high-performance AI/GenAI architectures featuring NVIDIA GPU clusters
- Design end-to-end AI Cloud and next-generation platforms optimized for deep learning workloads and distributed training.
- Architect HPC cluster topologies utilizing high-speed InfiniBand (NDR/HDR) and RoCE v2 interconnects for low-latency communication.
- Right-size platform components, including GPUs, CPUs, memory and NVMe storage for comprehensive client proposals.
- Architect distributed training and inference environments optimized for MPI frameworks and workload scheduling via Slurm.
- Desing scalable container orchestration platforms using Kubernetes and Kubeflow to manage AI workloads.
- Propose optimized inference strategies using vLLM, Triton, and TensorRT-LLM to meet specific latency and throughput KPIs.
- Should have experience on RAG systems and multi-agent orchestration frameworks like LangGraph and agentic ecosystems.
- Develop private AI cloud environments focused on data sovereignty and regulatory compliance, such as the India DPDP Act.
- Define integration strategies for LLMs and open-source models within existing enterprise data systems, APIs, and knowledge graphs.
- Establish reference architectures for CI/CD/CT pipelines and automated model retraining workflows to ensure reproducibility.
- Implement automation and observability frameworks for monitoring GPU utilization, performance tuning, and failure handling.
- Drive technical validation through Proof of Concept (PoC) engagements, focusing on scalability and performance benchmarks for LLM training.
- Establish Infrastructure-as-Code (IaC) practices to ensure reproducible and reliable cluster deployments.
- Collaborate with C-suite stakeholders and cross-functional teams to drive technical decision-making, innovation, and roadmap alignment.
📌 AI Architect (Chennai)
🏢 Larsen and Toubro (L&T)
📍 Chennai
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.