10 Aug
|
ACURA SOLUTIONS
|
Mumbai
10 Aug
ACURA SOLUTIONS
Mumbai
We are looking for a Sr.AI Scientist to head our Voice AI Center of Excellence (COE).
Our voice platform is already generating revenue in the real estate, higher education, and government mid-markets, leveraging a robust campaign management pipeline backed by
Vapi and local Tata SIP trunks.
Your mission is two-fold: First, optimize our current cloud-based agent behaviors to dominate the automated outbound market. Second, lead our technical migration toward a
100% in-house, on-premises voice stack. You will build the local STT SLM TTS infrastructure to serve data-sensitive enterprise clients, laying the foundation for our native,
end-to-end Audio-to-Audio (A2A) models.
Core Responsibilities
Orchestration Prompt Engineering: Maximize the performance, context-handling,
and conversational flow of our current Vapi-backed enterprise voice agents.
On-Premises Stack Development: Build and benchmark our localized, air-gapped voice alternative using high-velocity open-source components (e.g., Faster-Whisper for STT, highly optimized 3B/8B Small Language Models, and Kokoro or Qwen3-TTS for low-latency synthesis).
Streaming Telephony Integration:
Optimize bidirectional audio token streaming via WebSockets to keep glass-to-glass latency under 300ms when connected to local
SIP trunks.
A2A RD Roadmap: Begin foundational architecture design for native Audio-to-
Audio neural processing (leveraging concepts from neural audio codecs like EnCodec or architectures like Kyutais Moshi) to eliminate the traditional cascaded pipeline.
Required Technical Skillset Experience: 35 years of pure-play experience building real-time speech, NLP, or conversational AI systems.
Speech Technologies: Deep familiarity with advanced STT (Whisper-live, Parakeet
TDT) and up-to-date streaming TTS architectures (StyleTTS2, Fish Speech, CosyVoice2).
LLM Fine-Tuning: Proven experience fine-tuning and prompting open-source Small
Language Models (Llama-3-8B, Phi-3, Qwen) specifically for tool-calling and real-time structured dialogue.
Protocols Audio: Strong understanding of WebSockets, raw audio byte manipulation (16kHz, int16, mono), and real-time streaming constraints. Experience with VoIP/SIP trunks is a major plus.
This job is provided by Shine.com
📌 Sr.AI Scientist (Voice Lead) (Mumbai)
🏢 ACURA SOLUTIONS
📍 Mumbai