29 Aug
|
Simpplr
|
Gurugram
About the role
We are looking for a Lead Voice AI Engineer to build production-grade Voice Agents for frontline customer and employee support. You will lead the design of low-latency, real-time voice systems combining ASR, TTS, LLMs, conversational AI, enterprise workflows, knowledge retrieval, compliance, and human handoff.
This is a hands-on technical leadership role for someone who can take Voice AI from architecture to production.
Responsibilities
- Design and build the real-time voice runtime for live conversations.
- Build and optimize streaming ASR, TTS, VAD, endpointing, turn-taking, and barge-in.
- Develop Voice Agents that support multi-turn, multi-intent conversations, context switching, clarification, and recovery.
- Integrate Voice Agents with workflows, APIs, CRM, ITSM, knowledge bases, and enterprise systems.
- Build secure identity verification, consent, privacy, audit, and compliance controls.
- Implement warm transfer, callback, queue routing, and seamless human handoff with full conversation context.
- Optimize multilingual voice quality across accents, noisy environments, latency, and naturalness.
- Build evaluation frameworks for WER, intent accuracy, response latency, containment, resolution, escalation, and CSAT.
- Establish production observability across the full call path: ASR → LLM → tools → TTS.
- Evaluate and integrate leading speech, telephony, and AI technologies.
- Define architecture, engineering standards, and production readiness for the Voice AI platform.
- Mentor engineers and lead critical technical design reviews.
Minimum qualifications
- 7+ years of software engineering experience.
- Strong experience building production distributed or real-time systems.
- Hands-on experience with Conversational AI, Voice AI, Speech AI, or LLM-based agents.
- Solid programming skills in Python, Java, Go, or equivalent.
- Experience with APIs, streaming systems, asynchronous architectures, and cloud-native platforms.
- Strong understanding of system design, scalability, reliability, and observability.
Preferred qualifications
- Experience with Deepgram for real-time ASR and streaming speech recognition.
- Experience with LiveKit for WebRTC, real-time audio, voice-agent runtime, and session orchestration.
- Experience with ElevenLabs for low-latency, natural TTS and conversational voice experiences.
- Experience with OpenAI, Azure Speech, Google Speech, or similar ASR/TTS technologies.
- Experience with WebRTC, SIP, RTP, WebSockets, Twilio, or contact-center platforms.
- Experience with LLM agents, tool calling, RAG, LangGraph, or similar orchestration frameworks.
- Experience integrating enterprise systems such as Salesforce, ServiceNow, Jira, Zendesk, or Workday.
- Experience with multilingual speech, accent handling, noisy environments, PII redaction, and call-recording controls.
- Experience building high-scale, multi-tenant SaaS platforms.
What success looks like
- Voice Agents feel natural and responsive in real-time conversations.
- Users can interrupt naturally and change context without breaking the conversation.
- The system works reliably across languages, accents, and noisy environments.
- Voice Agents securely execute enterprise workflows and grounded knowledge retrieval.
- Complex cases escalate to humans with full context.
- The platform meets measurable targets for latency, accuracy, reliability, containment, resolution, and customer satisfaction.
Engineering principle
Voice is not chat with audio. Voice is a real-time interaction model with different requirements for latency, interruption, identity, compliance, failure handling, and human handoff.
📌 Lead Voice AI Engineer (Gurugram)
🏢 Simpplr
📍 Gurugram