You'll own voice AI at JoyzAI. All of it. The real-time pipeline (audio in, speech understood, LLM reasoning, tools called, speech out), the telephony that gets the call connected, the platform a client uses to configure and run their agent, and the webhooks and CRM plumbing that turn a call into a lead. It's a working product with real clients on it today. Your job is to take it from working to something we'd bet the company on, and to build the team that keeps it that way.
This is a head-of role at a company of under 15 people, which makes it a player-coach job.
You set the technical direction for voice: what we build on, what we buy, what we rip out. And you write a lot of the code yourself. When a call stutters, drops, or talks over the customer, you're the one who finds out why, fixes it at the root, and changes the system so it can't happen again. When a client wants an agent that does something we've never done, you decide whether the platform can, and then you make it so.
What you'll do
Own the voice stack
- Set the architecture and roadmap for the real-time calling pipeline: telephony ↔ audio streaming ↔ speech/real-time models ↔ tool calls ↔ speech out, with latency you can't feel.
- Own the hard voice problems, with opinions earned in production: turn-taking and barge-in, silence and voicemail detection, noise and echo, language switching mid-call (English ↔ Hinglish ↔ regional), and graceful recovery when a model or carrier hiccups.
- Make the build-vs-buy calls across STT/TTS, real-time LLMs, VAD, and telephony carriers. Run the evaluations, own the vendor relationships, decide what we build on next, and know when it's time to switch.
- Design non-blocking tool calling inside live calls (lookups, bookings, CRM writes) so the agent never goes silent or loses the thread.
- Define what "good" means for a voice agent: the quality bar, the evaluation framework, the test harness. We should know a change made agents better before a client tells us.
Build the platform around the calls
- Own the Node.js/Express services and APIs that have to hold up as call volume grows: call orchestration, queueing, webhooks, retries, recording and transcript storage.
- Own the product around the calls: agent configuration (tasks, behaviour, guardrails, AI config), call logs and transcripts, analytics, and bulk-calling campaign controls.
- Wire calls into the rest of the platform: CRM contacts and tickets, WhatsApp follow-ups after a call, calendar and booking integrations, Google Sheets and Meta lead-form ingestion.
- Keep the Angular frontends our team and clients use (the agent configurator, the inbox, the call review screens)
clean and usable, whether you're writing them or reviewing them.
Reliability, cost & scale
- Instrument everything: time-to-first-word, per-call latency breakdowns, drop rates, tool-call failures, model cost per minute. Own the dashboards that tell us something's wrong before a client does.
- Treat cost per minute as a number you're accountable for. Bring it down without hurting quality: model choice, streaming strategy, caching, carrier routing.
- Own production. Debug incidents on real client calls, reproduce them, fix them at the root, and build the incident and on-call practice so the next one gets caught faster.
Lead the team
- Build and lead the voice engineering team. Hire, set the bar, review the code, and make the engineers around you better.
- Own the engineering standard for voice: architecture decisions, code quality, documentation, what ships and what doesn't.
- Pair with the prompt engineer and customer success on what the platform needs for agents to behave well: new tool types, better fallbacks, a proper test harness for prompts.
- Be the technical authority in front of clients on the hard deployments, the ones where the answer isn't in the docs yet.
- Turn what you learn from client deployments into reusable platform capability, so the next voice agent takes hours to stand up, not days.
Roughly 60% of your time is hands-on: architecting, coding, debugging live calls, shipping. The other 40% is leading: the team, the roadmap, model and vendor decisions, and the client conversations where the platform's direction gets decided. That ratio shifts toward leading as the team grows, but you never stop being able to open the code and fix it yourself.
What we're looking for
Three things are non-negotiable. Please apply only if you meet all three. They're the first thing we screen for.
- You have successfully deployed voice AI agents to production, and owned them there. Not a demo, not a hackathon, not a pilot that quietly ended. A system where real customers spoke to an AI over a phone line or live audio stream, at real volume, for months, and it stayed in production because it worked. You owned the pipeline end to end: streaming audio, speech and real-time models, latency, interruptions, the failures at 2am,
the cost. In our first conversation we'll ask what you shipped, how many calls it handled, what broke, and what you did about it. Come with specifics. Experience at a Voice AI company is a robust signal.
- You can build the whole thing yourself. Backend services, APIs, real-time transport, databases, and a frontend a client can use without a walkthrough. You've done this as an individual contributor and you can still do it. A head of voice who can't debug the socket is not who we're hiring.
- You've set the technical direction for a voice product before. As a founding engineer, tech lead, architect, or head of engineering. The title matters less than the fact that people looked to you for what to build and how, and you were right often enough that the product worked.
Beyond that, we don't care much about degrees or years. We care that you can:
- Work fluently across the MEAN stack — MongoDB, Express, Angular, Node.js — with TypeScript throughout.
- Design and debug real-time systems: WebSockets, WebRTC, streaming audio, event-driven architectures, queues.
- Integrate with telephony/SIP providers and messaging APIs (WhatsApp Business API) and handle their quirks in production.
- Work with LLM APIs, tool/function calling, and speech models, and reason about latency, cost, and failure modes, not just accuracy.
- Design MongoDB data models and APIs for a multi-tenant SaaS: contacts, conversations, tickets, workflows, role-based access.
- Hire well, review well, and explain a hard technical call to a founder or a client in plain language.
- Measure and profile before you optimise, write code others can read, and document what you built.
Nice to have: deep experience with real-time voice frameworks or APIs (OpenAI Realtime, LiveKit, Pipecat, or similar); Indian telephony providers; Hinglish or Indic speech models; GCP/AWS at scale; having scaled a voice product from its first client to serious call volume; having built a voice team from scratch.
Full-time, on-site in Noida.
Why JoyzAI?
✨ Utmost freedom and autonomy — voice AI is yours. Make the decisions, set the direction, do it the way you think is best. We will only suggest.
✨ Complete responsibility. You are the taskmaster. Take responsibility for it.
✨ Top-of-the-market pay – we reward high performance with the best packages.
✨ 5-day working – weekends are yours to recharge.
✨ Work from office in Noida – collaborate, learn, and grow with a passionate team.
✨ High-growth startup – opportunity to see first hand how AI brings change and be a part of that.
📌 Head of Voice AI Development (Noida)
🏢 JoyzAI
📍 Noida