19 Aug
|
GyanSys
|
Bengaluru
Experience: 5–7 years
Role Overview
We are looking for an AI Inference Engineer with solid hands-on experience in LLM/GenAI model inference, optimization, deployment and GPU computing. The engineer will work across the AI inference stack, from model optimization and runtime development to scalable production serving.
The ideal candidate should have a strong understanding of Transformer architectures, GPU systems, inference runtimes and distributed computing, with experience optimizing AI workloads for performance, latency, throughput and cost.
Key Responsibilities
- Design, develop and optimize LLM/SLM and Generative AI inference pipelines for production workloads.
- Deploy and scale models using inference frameworks such as vLLM, SGLang, TensorRT-LLM or equivalent.
- Optimize latency, throughput, GPU utilization, memory footprint and inference cost.
- Work on KV-cache optimization, continuous batching, speculative decoding, quantization and model parallelism.
- Optimize inference across multi-GPU and multi-node environments.
- Analyze GPU compute and memory bottlenecks using profiling and benchmarking tools.
- Work with CUDA, NCCL, GPU memory management and NVIDIA GPU architectures.
- Evaluate and benchmark models across different GPU configurations and inference runtimes.
- Develop model serving solutions using Kubernetes, Docker and GPU orchestration platforms.
- Collaborate with ML engineers and researchers on model training, fine-tuning and inference optimization.
- Support SFT, LoRA/QLoRA and PEFT workflows and understand their impact on inference performance.
- Build automated performance benchmarking and evaluation frameworks for AI models.
- Implement monitoring and observability for production inference workloads.
- Work with MLOps/LLMOps teams on model versioning, deployment, CI/CD and production lifecycle management.
- Troubleshoot production issues involving GPU utilization, memory fragmentation, latency, throughput and scaling.
Required Skills
-
📌 AI Inference Engineer – LLM (Bengaluru)
🏢 GyanSys
📍 Bengaluru