05 Aug
|
HCL Technologies
|
Noida
05 Aug
HCL Technologies
Noida
Sr DomainTechnical Architect
Experience: 8 to 12 years
Location: Noida, India
Skills: Nvidia GPU Infrastructure, Kubernetes, GPU Cluster Administrator, Infrastructure SME, RCA, NVIDIA A100, NVIDIA H100, NVIDIA L40, AMD Instinct, TPU, CUDA, cuDNN, TensorRT, BeeGFS, Lustre, Ceph, InfiniBand, RoCE, RDMA, NVLink, NVIDIA GPU Operator, Kubeflow, Ray, MLflow, Slurm, PBS, NCCL, PyTorch Distributed, Horovod, DeepSpeed, vLLM, Triton Inference Server, Terraform, Helm, Customize, GitOps, ArgoCD, Flux, Prometheus, Grafana, NVIDIA DCGM, OpenTelemetry
Job Summary
Business Development Group, HCBU, HCLTech www.hcltech.com Digital Foundation / Full time We are HCLTech, one of the fastest-growing large tech companies in the world and home to 220,000+ people across 60 countries, supercharging progress through industry-leading capabilities centered around Digital, Engineering and Cloud. The driving force behind that work, our people, are diverse, creative, and passionate, raising the bar for excellence on a regular basis. We, in turn, work hard to bring out the best in them as we strive to help them find their spark and become the best version of themselves that they can be. If all this sounds like an environment you’ll thrive in, then you’re in the right place. Join us on our journey in advancing the technological world through innovation and creativity.
The Role
The AI Infrastructure Engineer (L3) provides advanced engineering and architectural expertise for high‑performance AI and ML infrastructure. This role focuses on building, optimizing, and scaling GPU/accelerator environments and distributed systems for large‑scale training and inference workloads. Competency Focus: High‑performance computing (HPC), distributed systems, Kubernetes, GPU orchestration, cloud optimization
Responsibilities:
• Deploy, configure, and manage GPU and AI accelerator platforms (NVIDIA A100/H100/L40, AMD Instinct, TPU).
• Troubleshot GPU hardware and software issues, including failures, thermal throttling, PCIe/NVLink topology, and driver conflicts.
• Install, upgrade, and maintain GPU software stacks, including drivers, CUDA, cuDNN, TensorRT, and firmware.
• Perform capacity planning and resource optimization for AI training, fine‑tuning, and inference workloads.
• Optimize Linux systems (Ubuntu, RHEL, Rocky) for AI/HPC workloads through NUMA, kernel, and clock tuning.
• Manage distributed and high‑performance storage systems, including BeeGFS, Lustre, Ceph,
Key Responsibilities
• Operate high‑bandwidth, low‑latency networks, including InfiniBand, RoCE, RDMA, and NVLink.
• Administer Kubernetes GPU clusters, leveraging NVIDIA GPU Operator, device plugins, MIG, and node feature discovery.
• Support AI and HPC orchestration platforms, including Kubeflow, Ray, MLflow, and Slurm/PBS.
• Configure and manage GPU scheduling and sharing strategies, such as node pools, quotas, job queues, and fair‑share policies.
• Optimize distributed training workflows using NCCL, PyTorch Distributed, Horovod, and DeepSpeed.
• Operate and tune LLM and inference runtimes, including vLLM, Triton Inference Server, and TensorRT‑LLM.
• Monitor and tune GPU utilization, memory allocation, and container-level performance.
• Automate cluster provisioning and operations using Terraform, Helm, Customize, and GitOps (ArgoCD/Flux).
• Build automation for GPU diagnostics, node onboarding, and model deployment workflows.
• Implement observability and telemetry using Prometheus, Grafana, NVIDIA DCGM, and OpenTelemetry.
• Lead deep‑dive root cause analysis for GPU, network, storage, and orchestration issues.
• Provide L3 support and work with L2/L1 teams for escalations.
• Drive production readiness, patching, hotfix rollout, and reliability improvements across AI infrastructure.
• Troubleshoot & escalation for complex platform failures Deep debugging of: NCCL hangs,
GPU fabric issues and co-ordinate with OEM and support vendors on critical issues.
• Review RCA, architecture documents, and change plans.
• Act as technical advisor to leadership and customers.
Qualifications & Experience
• Bachelor’s degree in computer science, Engineering, Information Technology, or related field.
• 8–12 years of overall infrastructure or platform engineering experience.
• 4–6 years of specialized experience supporting AI/ML workloads.
• Demonstrated experience in large‑scale GPU/accelerated computing and distributed systems.
• Strong experience in Kubernetes, containerization, and orchestration tools.
• Understanding of AI workload and MLOps.
Certifications Required
• NVIDIA Certified Associate – AI Infrastructure.
• NVIDIA NPN Certification.
• NVIDIA Base Command Manager certification.
• AWS Solutions Architect Associate.
• CKA – Certified Kubernetes Administrator.
• CKAD – Certified Kubernetes Application Developer.
How You’ll Grow
At HCLTech, the growth of an L3 AI Infrastructure Engineer is closely aligned with the organization’s competency framework, which emphasizes technical excellence, collaboration, continuous learning, and Ideapreneurship. At this level, engineers expand their impact through cross‑team collaboration, working with application, data science, cloud, SRE, security, and FinOps teams to design, operate, and optimize scalable AI platforms that directly support customer and business requirements. Regular vendor interaction further accelerates growth, as L3 engineers engage with OEMs and technology partners to resolve critical platform issues, lead joint root cause analyses, and contribute to roadmap and early‑access discussions, strengthening HCLTech ecosystem value and delivery confidence. Sponsored certifications play a key role in enhancing future‑ready skills, enabling engineers to deepen expertise in GPU platforms, Kubernetes, cloud‑native AI, and HPC technologies while applying this knowledge to live customer environments.
📌 Sr DomainTechnical Architect (Noida)
🏢 HCL Technologies
📍 Noida