10 Sep
|
Mobiloitte
|
Okhla
Job Description
AI Engineer — Agentic AI, RAG & Low-Latency Voice Systems
n
Mobiloitte Technologies (I) Pvt. Ltd. | New Delhi (on-site / hybrid) | Full-time | 5+ years
n
Compensation: No ceiling for the right candidate.
n
n
Two things decide this role: how well you handle latency, and how far past basic RAG you can build.
n
n
Mobiloitte builds production AI for enterprise clients across India, the US, UK, UAE, Singapore and South Africa. We're hiring a senior AI Engineer to own our agentic AI, retrieval and voice stack — systems that go live, carry real call volume, and stay up.
n
n
1. LATENCY IS THE JOB, NOT A DETAIL
n
n
Our voice bots handle inbound and outbound calls where a 400ms difference decides whether the conversation feels human. We need someone who treats latency as an engineering discipline:
n
n
- Owns end-to-end voice latency from user speech to first token of bot audio — and can quote the numbers they achieved
n
- Profiles and attacks each hop: STT, retrieval, inference, TTS, network
n
- Streaming responses, speculative/partial generation, response chunking to shorten time-to-first-audio
n
- Barge-in and interruption handling that actually works mid-sentence
n
- Model caching, warm pools, batching, concurrency control under real call load
n
- Knows when the fix is architectural, not a bigger GPU
n
n
If you can't tell us what your p95 latency was and what you did to bring it down, this isn't the role.
n
n
2. AGENTIC AI — BEYOND BASIC RAG
n
n
Single-shot retrieve-and-answer is table stakes. We're building systems that plan, act and self-correct:
n
n
- Agent orchestration — LangGraph, CrewAI, AutoGen, OpenAI Agents SDK, or your own framework (we care that you've built it, not which one)
n
- Tool and function calling, API/DB actions, structured outputs
n
- Multi-step reasoning, task decomposition, planning loops with retry and fallback
n
- Multi-agent patterns — routing, supervisor/worker, handoffs
n
- Conversation and long-term memory management across sessions
n
- Advanced retrieval: agentic and query-rewriting RAG, hybrid search, reranking, graph/multi-hop retrieval, chunking strategy chosen deliberately
n
- Evaluation and guardrails — hallucination reduction, groundedness scoring, tracing and observability (LangSmith, Langfuse or similar)
n
- Cost control: token budgeting, context-window discipline, model routing
n
n
3. VOICE — INBOUND AND OUTBOUND
n
n
- Full pipeline: STT → LLM/agent/RAG → TTS
n
- Whisper / Faster-Whisper / Deepgram; Piper / Coqui / XTTS or equivalent
n
- SIP and telephony integration, WebSockets, real-time streaming
n
- Built the stack rather than wrapping Vapi, Retell or Bland end to end
n
n
4. PRIVATE / SELF-HOSTED DEPLOYMENT
n
n
Many of our clients cannot send data to public APIs. You should be able to:
n
n
- Deploy Llama, Qwen or Mistral inside a client VPC or on-prem
n
- Run vLLM, Ollama or TensorRT-LLM in production
n
- Size GPUs and VRAM, apply INT8/INT4 quantization, tune inference
n
- Handle concurrency, autoscaling, secure API exposure, auth and data isolation
n
- Explain how you'd replace an OpenAI dependency with a self-hosted architecture — this earns robust preference
n
n
ON COMPENSATION
n
n
We've deliberately not published a band. For an engineer who has genuinely built agentic systems, driven latency down in production, and run models on their own infrastructure, compensation is not the constraint — we'll match what the work is worth.
n
n
Be honest with yourself before applying: if your AI work is mainly calling third-party APIs and writing prompts, we'll find that out in the first twenty minutes. If you've built agents, fought latency, and put a model on your own GPU, we want to talk.
n
n
HOW TO APPLY
n
n
Apply here and include one thing: a link, repo, demo video or short architecture note for a system you personally built — and tell us what you did on it.
n
n
Thanks
n
Team-HR
📌 AI Engineer -- Agentic AI, RAG & Low-Latency Voice Systems | Okhla | 5+ yrs
🏢 Mobiloitte
📍 Okhla