29 Aug
|
VoiceCare AI
|
Bengaluru
29 Aug
VoiceCare AI
Bengaluru
About the role You'll build and optimize the AI systems behind our platform - starting with the real-time conversational engine behind Joy, and extending across the surfaces that surround it: call summarization and analytics, multi-modal orchestration, document intelligence, and computer-use agents. This is a hands-on engineering role at the intersection of applied LLMs, real-time audio, and agentic workflows - where a 200 ms improvement in time-to-first-token is a real, measurable win, and where a document pipeline or a browser agent has to be just as reliable in a clinical, PHI-sensitive setting.
What you'll do
- Build and optimize Joy's real-time voice pipeline (streaming STT → LLM → TTS) on a voice AI orchestration framework, targeting sub-second p95 time-to-first-token.
- Implement and tune latency techniques: sentence-chunked streaming TTS, speculative tool prefetching, filler speech, parallel tool execution, and prompt caching.
- Integrate LLMs for conversational turns - function/tool calling, prompt engineering, structured outputs, and grounded responses.
- Develop and improve the RAG layer over vector search: retrieval quality, knowledge-base grounding, and refusal behavior on out-of-scope or clinical questions.
- Benchmark models across multiple frontier LLM providers on latency, cost, and quality - and turn results into concrete configuration and routing decisions.
- Build and maintain evaluation harnesses covering retrieval metrics and generation metrics (faithfulness, correct-refusal rate, containment).
- Integrate telephony and messaging (SIP trunking / SMS via a programmable telephony provider) and coordinate with EHR middleware for appointment and advantages workflows.
- Build call summarization that turns raw transcripts into structured post-call notes and extracted fields, and feed the analytics layer that surfaces KPIs like containment rate, resolution, and call quality.
- Contribute to the orchestration layer that coordinates AI across modalities (voice, SMS, document) - session/state management, routing between agents, and consistent tool-calling across channels.
- Build the document intelligence pipeline for healthcare benefits-verification documents - schema-first extraction and summarization with a two-tier model cascade and confidence scoring using layout-aware document-parsing libraries, batch processing at scale.
- Develop computer-use / browser agents that operate EHR portals and web workflows (navigation, form-filling, data retrieval) reliably and safely.
- Deploy, monitor, and tune services on a serverless container platform; add the observability needed to debug real calls and agent runs.
- Partner closely with the backend and product teams to ship reliable, testable features.
- Own the AI/ML technical vision and roadmap across the full agent portfolio - voice, multi-modal orchestration, summarization/analytics, document intelligence, and computer-use agents - tied to clear business outcomes (containment, accuracy, cost per interaction).
- Lead, mentor, and grow the AI engineering team - hiring, leveling, career development, and day-to-day technical direction.
- Set the end-to-end architecture for the voice pipeline (STT → LLM → TTS, retrieval, tool calling, telephony) and for the orchestration layer that coordinates AI across modalities, document processing, and agentic web/EHR automation - and define the standards the team builds against.
- Define and drive service-level objectives per surface - e.g., p95 time-to-first-token for voice, grounding accuracy and ungrounded-clinical-claim rate for RAG, extraction accuracy and confidence thresholds for documents, task success rate for computer-use agents - and hold the team to them.
- Own model strategy: build-vs-buy decisions, provider and region selection, and a rigorous benchmarking practice across latency, cost,
and quality.
- Establish the evaluation and quality framework spanning retrieval, generation, summarization, document extraction, and agent task completion - plus the safety posture for clinical and grounded responses.
What we're looking for
- 2–5 years building production software, with meaningful applied LLM / ML work.
- Strong Python, including async and modern async web frameworks.
- Hands-on experience with LLM applications: prompting, tool/function calling, RAG, embeddings, and vector databases.
- Solid grasp of real-time or streaming systems and the discipline to reason about latency budgets.
- Comfort with cloud infrastructure (GCP) and containerized deployment.
- Agentic / computer-use experience - browser or desktop automation, tool-using agents, multi-step task orchestration
- Streaming STT / TTS systems.
- LLM / RAG evaluation tooling and frameworks.
- A data-driven, benchmark-first instinct - you'd rather measure than trust a vendor claim.
Nice to have
- Voice / telephony experience: SIP, G.711/G.722 codecs, VAD, and voice AI orchestration frameworks.
- Document intelligence / PDF extraction experience (layout-aware parsing, structured extraction).
- Experience in healthcare or another regulated, PHI-sensitive domain (HIPAA awareness).
---
About Voicecare VoiceCare AI is a Healthcare Administration General Intelligence company focused on revenue cycle management (RCM) and back-office operations for healthcare organizations. The company’s mission is to improve access, adherence, and outcomes for patients and healthcare teams through generative AI. Its enterprise-grade Agentic AI platform automates routine phone conversations and administrative tasks such as benefit verification, prior authorization, claims management, and prescription support.
VoiceCare AI is built with security and compliance at its core and is SOC 2 Type II attested and HIPAA compliant. Applicants joining VoiceCare AI will contribute to transforming healthcare operations with secure, scalable AI solutions.
📌 AI Engineer (Bengaluru)
🏢 VoiceCare AI
📍 Bengaluru