19 Aug
|
ixigo
|
New Delhi
Job Description
Voice agents fail in ways traditional software doesn't. ASR confidence drops on an accent and a tool call misfires. Latency breaks turn-taking and the LLM hallucinates a policy. A model swap silently regresses production and nobody catches it for a week.
We're building self-healing voice agents for enterprise customer support. This role owns the intelligence layer: the evals that catch failures before shipping, the observability that traces them across the pipeline, and the feedback loops that let agents fix themselves
What you'll own
• Evaluation infrastructure. Audio-native metrics for barge-in, prosody, and turn-taking. Adversarial datasets across accents and edge cases. LLM-as-judge rubrics for task success, tool-use correctness, and recovery.
• Observability across the pipeline. Tracing that correlates audio, STT, LLM reasoning, tool calls, and TTS to a single conversation. Analysis and alerting that surfaces cascade failures instead of hiding them.
• Self-improvement systems. Mine production traces for failure patterns, generate targeted training or prompt data, validate fixes with adversarial replay, and guardrail against regressions.
📌 Research Engineer - Agent Intelligence & Evaluation (New Delhi)
🏢 ixigo
📍 New Delhi