17 Sep
|
Elabs Infotech
|
India
17 Sep
Elabs Infotech
India
Experience: 512 Years
Location: Mumbai, Pune, Bangalore, Hyderabad, Chennai, Noida
Employment Type: Full time
Job Summary
We are looking for an experienced NVIDIA GPU Platform Engineer to build, operate, and optimize GPU-accelerated AI/ML infrastructure on Kubernetes and AWS. The role focuses on NVIDIA GPU platforms, LLM inference/model serving, GPU scheduling and partitioning, autoscaling, telemetry, and performance optimization.
Key Responsibilities
Design and manage NVIDIA GPU platforms for AI/ML workloads.
Deploy and manage CUDA, NVIDIA GPU Operator and DCGM.
Build and operate Kubernetes-based GPU workloads using GPU Device Plugins.
Implement GPU scheduling using Run or equivalent technologies.
Configure MIG and fractional GPU capabilities for efficient GPU utilization.
Deploy and optimize model-serving platforms including NVIDIA NIM, Dynamo, Triton, TensorRT-LLM and vLLM.
Implement KEDA-based autoscaling for inference workloads.
Develop and integrate OpenAI-compatible inference APIs.
Implement GPU telemetry, monitoring, utilization and performance dashboards.
Support LLM and RAG inference workloads, including vector database integrations.
Perform inference benchmarking, capacity planning,
performance tuning and optimization.
Work with NVIDIA AI Enterprise (NVAIE) on AWS/EKS and GPU-enabled cloud/edge platforms.
Support AWS Outposts G7 / Edge GPU platforms where required.
Mandatory Skills
Solid hands-on experience with NVIDIA GPUs and CUDA
NVIDIA GPU Operator, DCGM
Kubernetes GPU workloads and Device Plugins
Model serving: NIM / Triton / TensorRT-LLM / vLLM / Dynamo
GPU scheduling Run or equivalent
MIG / Fractional GPU
KEDA autoscaling
OpenAI-compatible inference APIs
GPU telemetry, monitoring and performance optimization
Good-to-Have Skills
NVAIE on AWS/EKS
AWS Outposts / Edge GPU platforms
LLM and RAG serving
Vector databases
GPU/inference performance benchmarking
TensorRT and CUDA performance tuning
AI/ML platform engineering experience
Ideal Candidate Candidates should have robust hands-on experience in GPU infrastructure + Kubernetes + AI model serving, rather than purely application-level ML development. Experience deploying and optimizing production-grade GPU inference platforms will be highly preferred.
📌 Nvidia Gpu Platform Engineer Bengaluru (India)
🏢 Elabs Infotech
📍 India