We are looking for an experienced Inference Engineer to build and optimize high-performance inference systems for Large Language Models (LLMs), Speech AI systems, and multimodal AI workloads.You will work on deploying production-grade AI systems with a robust focus on:
low latency,
high throughput,
GPU efficiency,
scalable serving infrastructure,
distributed inference,
and cost optimization.
This role sits at the intersection of:
systems engineering,
deep learning infrastructure,
distributed computing,
and production AI deployment.
You will collaborate closely with:
ML researchers,
platform engineers,
speech AI teams,
and product engineering teams.
📌 Inference Engineer – Llm & Speech Ai Bengaluru (India)
🏢 soket.ai
📍 India
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.