03 Aug
|
programming.com
|
India
03 Aug
programming.com
India
AI/ML Engineer – Model Optimization & Acceleration (8–10 Years)
- Location: Bengaluru, India
- Experience: 8–10 Years ( If you have experience more than 8 years only apply then)
- Open Positions: 3
We are looking for an experienced AI/ML Engineer to optimize and deploy machine learning models across heterogeneous platforms (CPU, GPU, and NPU). If you're passionate about building high-performance, production-ready AI systems and working on cutting-edge technologies, we'd love to hear from you!
Key Responsibilities
- Optimize AI models including LLMs, Diffusion Models, CNNs, Computer Vision, Multi-modal, and Speech Models.
- Port models across frameworks (PyTorch → ONNX → Runtime).
- Deploy and optimize models on GPU/NPU hardware accelerators.
- Improve inference latency, throughput, and memory efficiency.
- Implement quantization, model compression, and performance tuning.
- Profile,
benchmark, and debug AI system performance.
Required Skills
- Strong expertise in PyTorch and ONNX
- Proficiency in Python and C++
- Experience with CUDA, ROCm, or GPU acceleration
- Solid understanding of Transformers, CNNs, Deep Learning
- Hands-on experience in Model Optimization, Quantization, Inference Optimization, and Performance Tuning
Good to Have
- Edge AI / Embedded AI deployment
- Generative AI or Multi-modal AI
- Distributed inference or streaming pipelines
- TensorRT, OpenVINO (preferred)
Pay: ₹2,500,000.00 - ₹3,000,000.00 per year
Experience:
- AI/ML Engineer – Model Optimization & Acceleration: 8 years (Preferred)
Work Location: In person
📌 AI/ML Engineer – Model Optimization & Acceleration (India)
🏢 programming.com
📍 India