16 Sep
|
Elabs Infotech
|
Bengaluru
16 Sep
Elabs Infotech
Bengaluru
Experience: 512 Years
Location: Mumbai, Pune, Bangalore, Hyderabad, Chennai, Noida
Employment Type: Full-time
Job Summary
We are looking for an experienced NVIDIA GPU Platform Engineer to build, operate, and optimize GPU-accelerated AI/ML infrastructure on Kubernetes and AWS. The role focuses on NVIDIA GPU platforms, LLM inference/model serving, GPU scheduling and partitioning, autoscaling, telemetry, and performance optimization.
Key Responsibilities
- Design and manage NVIDIA GPU platforms for AI/ML workloads.
- Deploy and manage CUDA, NVIDIA GPU Operator and DCGM.
- Build and operate Kubernetes-based GPU workloads using GPU Device Plugins.
- Implement GPU scheduling using Run or equivalent technologies.
- Configure MIG and fractional GPU capabilities for efficient GPU utilization.
- Deploy and optimize model-serving platforms including NVIDIA NIM, Dynamo, Triton, TensorRT-LLM and vLLM.
- Implement KEDA-based autoscaling for inference workloads.
- Develop and integrate OpenAI-compatible inference APIs.
- Implement GPU telemetry, monitoring, utilization and performance dashboards.
- Support LLM and RAG inference workloads, including vector database integrations.
- Perform inference benchmarking, capacity planning,
performance tuning and optimization.
- Work with NVIDIA AI Enterprise (NVAIE) on AWS/EKS and GPU-enabled cloud/edge platforms.
- Support AWS Outposts G7 / Edge GPU platforms where required.
Mandatory Skills
- Strong hands-on experience with NVIDIA GPUs and CUDA
- NVIDIA GPU Operator, DCGM
- Kubernetes GPU workloads and Device Plugins
- Model serving: NIM / Triton / TensorRT-LLM / vLLM / Dynamo
- GPU scheduling Run or equivalent
- MIG / Fractional GPU
- KEDA autoscaling
- OpenAI-compatible inference APIs
- GPU telemetry, monitoring and performance optimization
Good-to-Have Skills
- NVAIE on AWS/EKS
- AWS Outposts / Edge GPU platforms
- LLM and RAG serving
- Vector databases
- GPU/inference performance benchmarking
- TensorRT and CUDA performance tuning
- AI/ML platform engineering experience
Ideal Candidate Candidates should have robust hands-on experience in GPU infrastructure + Kubernetes + AI model serving, rather than purely application-level ML development. Experience deploying and optimizing production-grade GPU inference platforms will be highly preferred.
📌 Nvidia GPU Platform Engineer (Bengaluru)
🏢 Elabs Infotech
📍 Bengaluru