06 Aug
|
Birdeye
|
Gurugram
Job Type Full-time Description About the Role We are looking for a Senior Voice AI Engineer to build and scale next-generation conversational Voice AI systems. You will work on real-time, low-latency voice pipelines involving Speech-to-Text (STT), Large Language Models (LLMs), Text-to-Speech (TTS), WebRTC, telephony, and streaming infrastructure. You should have hands-on experience building production-grade Voice AI agents using frameworks like Pipecat, LiveKit, or similar real-time communication platforms.
This role requires deep understanding of streaming architectures, WebSockets, voice quality optimization, interruption handling, and scalable backend systems.
About Birdeye
Birdeye is the leading agentic marketing platform for multi-location brands. Companies like H&R; Block, Aspen Dental, and Caesars Entertainment use Birdeye to manage marketing across thousands of locations — from how they get found, to how they convert, to how they retain customers. Our platform replaces disconnected point tools with AI agents that execute work at the location level — responding to reviews, updating listings, publishing content, and driving conversions.
Backed by Marc Benioff, Jerry Yang, and Accel-KKR, Birdeye was named to G2’s 2026 Best Agentic AI Products list — appearing alongside the world’s leading AI companies. We’re expanding rapidly into enterprise, with growing adoption across large, multi-location brands.
Responsibilities: Design and build production-grade Voice AI applications.
Develop real-time streaming pipelines for voice conversations.
Integrate STT, LLM, and TTS providers into low-latency conversational systems.
Build robust WebSocket/WebRTC infrastructure for bi-directional audio streaming.
Implement interruption handling (barge-in), turn detection, and Voice Activity Detection (VAD).
Optimize latency across the complete voice pipeline.
Integrate telephony providers such as Twilio, SIP, or LiveKit Telephony.
Develop scalable backend services capable of supporting thousands of concurrent voice sessions.
Design session management, state management, and conversation memory.
Build monitoring, observability, and analytics for voice conversations.
Deploy and operate Voice AI infrastructure on Kubernetes and cloud platforms.
Required Skills: Voice AI Solid understanding of conversational Voice AI systems
Experience building real-time AI voice assistants
Knowledge of latency optimization techniques
Understanding of conversational memory and dialogue management Speech Technologies Experience with OpenAI Realtime API, ElevenLabs, Deepgram, AssemblyAI, Google Speech, Azure Speech, Amazon Transcribe, Cartesia, or PlayHT
Streaming STT and TTS
Voice cloning and adaptive speech
Speaker diarization
Custom vocabulary and pronunciation dictionaries Real-Time Communication Pipecat
LiveKit
WebRTC
RTP
SIP
WebSockets
Server-Sent Events (SSE) Backend Development Python (preferred), FastAPI, AsyncIO
WebSocket servers, gRPC, REST APIs
Event-driven architecture
Redis, Kafka, RabbitMQ
PostgreSQL, MongoDB, ClickHouse (good to have) AI & LLM Experience with OpenAI, Anthropic, Google Gemini, or open-source LLMs
Prompt engineering, tool calling, function calling
RAG, agentic workflows, multi-agent orchestration, conversation memory Cloud & Infrastructure AWS, GCP, or Azure
Docker
Kubernetes
NGINX
Load Balancers
Autoscaling
CI/CD Requirements Preferred Experience: Built production Voice AI products
Experience with Pipecat in production
Experience with LiveKit Cloud or self-hosted LiveKit
Experience with Twilio Voice or SIP infrastructure
AI phone agents and call center automation
Noise suppression (Krisp, DeepFilterNet, RNNoise)
Experience supporting 1,000+ concurrent voice sessions
Multilingual Voice AI Qualifications: Bachelor's or Master's degree in Computer Science or related field.
6-8 years of software engineering experience.
3+ years building real-time streaming systems.
2+ years working on Voice AI or conversational AI products.
Strong system design and distributed systems experience. Why You'll Join Us: At Birdeye, we are relentless innovators driven by a singular goal: to lead our category with unparalleled excellence. We don't just set goals – we surpass them.
We're a team of doers who roll up our sleeves and get the job done, delivering on our promises with unwavering dedication. Working here means embracing a culture of action and accountability, where every person is empowered to make an impact. We don't just talk about making a difference – we make it happen.
📌 Senior Voice AI Engineer (Gurugram)
🏢 Birdeye
📍 Gurugram