About The Role
Build phone-native AI. You will engineer the streaming pipeline, the turn-taking and the telephony that make a voice agent feel like a conversation instead of an IVR.
What You'll Do
Build streaming speech pipelines: recognition, understanding, synthesis
Engineer for sub-second end-to-end response latency
Handle barge-in, turn-end prediction and interruption gracefully
Integrate telephony and handle real-world network conditions
Build multilingual conversation support with per-language evaluation
Wire booking and CRM actions as gated tools
What You'll Bring
2+ years backend or real-time systems engineering
Python or Node.js and comfort with streaming and websockets
Understanding of audio processing basics
LLM experience with latency as a first-class constraint
Patience for debugging problems that only appear on real calls
Nice to Have
Telephony experience: Twilio, SIP or WebRTC
Speech recognition or TTS integration experience
Multilingual or Indic language speech experience
What You'd Build
These aren't hypothetical projects, they're live products you can try before your first interview.
The StackBinary MarTech Suite, The full product line, every system we ship, most with live demos.
Why Join StackBinary™?
Adaptable working hours
Remote-friendly culture
Learning & development budget
High-ownership projects
Pragmatic engineering culture
Work with cutting-edge tech
Ready to Apply?
Join our team of builders who love shipping quality software.
We connect with shortlisted candidates through LinkedIn or our official email IDs.
Follow StackBinary on LinkedIn →
Questions about this role?
[email protected]
📌 Conversational AI Engineer (Mumbai)
🏢 Stackbinary
📍 Mumbai