Voice AI Expert to Make Our Voice Bot Faster & More Human (Fish Audio + Kokoro) (India)

Voice AI Expert to Make Our Voice Bot Faster & More Human (Fish Audio + Kokoro) (India)

29 Aug
|
BizFinder.ai
|
India

29 Aug

BizFinder.ai

India

We run a product with a real-time AI voice — It works and handles real calls today. However

We need an expert voice-AI engineer to obsess over two things: making it sound genuinely human, and making it respond instantly. Everything else is secondary.

The mission: humanness + speed

Make it human. Callers should forget they're talking to a machine — natural prosody, warmth, and pacing from our TTS; clean barge-in/interruption handling and endpointing so it never talks over people or leaves dead air; natural confirmation of spelled-out emails and phone numbers; no stiff "AI" phrasing.

Make it fast. We're around ~5s voice-to-voice and want to get under 2s without hurting naturalness. We have a tuning plan (VAD/endpointing, LLM prefix caching + speculative decoding, STT decode params, TTS warm-up/streaming, pipeline parallelism) — validate, execute, and push it further. Every 500ms cut makes it feel more alive.

Our current TTS — this is what you'll tune

We are standardizing on two TTS engines and want an expert to get maximum humanness and minimum latency out of these specifically (not to swap in new ones):

Local / self-hosted Kokoro TTS on our own GPU (NVIDIA A6000) — low-latency, no per-minute cost.

Fish Audio TTS (s2.1-pro) — our most expressive, human-sounding voice.

Direct, hands-on Fish Audio and/or Kokoro experience is a major plus — knowing their params, streaming behavior, warm-up, and reference voices.

Rest of the stack

Pipecat + Daily WebRTC → Whisper STT + Qwen via vLLM → local Kokoro.

LiveKit Agents → Inworld STT + LLM → Fish Audio or Kokoro (per call).





Python · asyncio · aiohttp · Docker · Postgres (call/turn metrics).

Must-have experience

Shipped real-time, low-latency voice agents to production — made one sound human AND respond rapid, and can prove it with numbers.

Deep knowledge of the STT → LLM → TTS streaming pipeline and where both the milliseconds and the naturalness come from.

Hands-on with Fish Audio and/or Kokoro (or very close analogues), plus Pipecat, LiveKit Agents, Daily, WebRTC, Whisper/faster-whisper, vLLM.

Strong Python + asyncio; comfortable operating self-hosted GPU inference (CUDA/VRAM on an A6000).

You measure both speed and human-ness — p50/p95 latency, TTFB budgets, and structured naturalness scoring.

Nice to have

Prosody/expressiveness & voice-clone/reference-audio tuning · receptionist persona prompt engineering · barge-in/interruption UX · frustration/sentiment detection & human escalation · staged production rollouts.

Please answer when you apply

Describe one time you made a voice agent sound more human or respond faster — before/after numbers and how you measured it.

Your hands-on experience with Fish Audio and/or Kokoro (or the closest engines you've shipped), and what you tuned to improve naturalness and TTFB.

In 2–3 sentences: how would you get a ~5s voice-to-voice loop under 2s while making it sound more human, not less?

Your availability, hourly rate, and time-zone overlap with US hours.

- Tip: We read applications that engage with the actual problem far more closely than generic ones. Please skip the templated cover letter — talk concretely about prosody, latency budgets, and Fish/Kokoro tuning.

📌 Voice AI Expert to Make Our Voice Bot Faster & More Human (Fish Audio + Kokoro) (India)
🏢 BizFinder.ai
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: voice ai expert to make our voice bot faster & more human (fish audio + kokoro) (india) / india