We Have 2 different requirement below is the JD's for both the requirement:
ROLE-1 : AI Infrastructure Engineer (GPU Computation)
ROLE 2:AI Software Engineer (Generative AI)
1. AI Infrastructure Engineer (GPU Computation)
Experience Required: 5+ years in AI/ML infrastructure or high-performance computing
Location: Delhi NCR
Infrastructure Environment
Hybrid both cloud (e.g. AWS) and on-premises GPU clusters
Key Responsibilities
- Design, configure, and optimize GPU compute pipelines on-premises and cloud-based to hit target utilization levels (~90%).
- Profile training and inference workloads to identify and remove hardware-level bottlenecks, regardless of where they're hosted.
- Recommend and implement scaling, batching, and parallelization strategies across on-prem and cloud GPU clusters.
- Decide, workload by workload, whether on-premises or cloud compute is the better fit, and manage the two consistently.
- Collaborate continuously with the AI Software Engineer to optimize model performance, compute utilization, deployment architecture, and experimentation speed.
- Set up monitoring and reporting on GPU utilization, training throughput, and compute efficiency metrics across both environments.
- Identify opportunities to improve platform infrastructure capacity planning, cost efficiency, reliability ahead of being asked.
- Support a hypothesis-driven experimentation culture by enabling quick, low-friction infrastructure for rapid prototyping and iteration.
Success Metrics
- GPU utilization against target (~90%).
- Training/inference throughput and experiment turnaround time.
- Infrastructure reliability (uptime, incident frequency).
- Cost optimization across cloud and on-premises spend.
Requirements
- 5+ years of hands-on experience in AI/ML infrastructure or high-performance computing.
- Strong expertise in GPU architecture and tooling (e.g., NVIDIA CUDA, TensorRT).
- Experience with distributed training/inference frameworks (e.g., DeepSpeed, Horovod,