Role Overview We are seeking a highly experienced Lead AI Engineer to architect, deploy, and scale intelligent, agentic voice systems capable of handling massive production traffic. This role focuses on building ultra-low-latency, full-duplex conversational voice agents using cascaded architectures (ASR -> LLM -> TTS) tailored for Indian languages.
You will act as the technical bridge between complex client workflows and scalable AI capabilities, ensuring robust backend execution, optimal resource efficiency, and seamless human-computer interaction at enterprise scale.
Core Responsibilities
- Full-Duplex Voice Architecture: Design and orchestrate conversational AI pipelines without relying on native multimodal/full-duplex LLMs. Build and tune highly responsive turn-taking logic, Voice Activity Detection (VAD), and barge-in/interruption handling across cascaded ASR, LLM, and TTS components.
- Telephony & API Integration: Seamlessly bridge AI inference pipelines with standard telephony APIs (Twilio, Plivo, Exotel) for inbound and outbound agent call flows, managing basic call state (transfers, hold, drop detection).
- Agentic Workflow Orchestration: Collaborate directly with clients to deconstruct complex business requirements and operational workflows. Translate these into deterministic agentic capabilities, utilizing state machines, tool-calling, and external API integrations to execute multi-step reasoning tasks.
- Latency & Efficiency Optimization: Drive hardcore performance tuning across the entire stack. Optimize Time-To-First-Token (TTFT), Time-To-First-Audio (TTFA), and Real-Time Factor (RTF) over phone lines.
Implement model quantization, KV cache optimization, dynamic batching, and effective model serving to minimize latency under heavy concurrent loads.
- Indic Language Mastery: Lead the development of multilingual systems that natively handle the phonetic and linguistic nuances of Indian languages. Solve complex challenges related to code-switching (e.g., Hindi-English, Telugu-English), regional accents, and low-resource language modeling.
- Backend & Production Scale: Architect resilient, event-driven backend systems capable of sustaining high-throughput production traffic. Manage stateful asynchronous processes, distributed microservices, and robust data pipelines to ensure zero-downtime deployments and real-time observability.
Required Qualifications & Experience
- Experience Baseline: 8+ years of overall software engineering and AI/ML experience, with a strict minimum of 3+ years directly architecting and deploying agentic LLM systems and complex conversational AI in production.
- Production System Expertise: Deep understanding of backend engineering for high-concurrency environments. Proven experience with distributed systems, event-driven architectures (e.g., Apache Kafka), workflow orchestrators (e.g., Temporal), and high-performance databases (e.g.,
PostgreSQL, ClickHouse).
- Conversational AI Depth: Strong operational knowledge of speech processing models (ASR/TTS) and streaming protocols (WebRTC, gRPC, WebSockets). You must know how to handle endpointing, stream buffering, and state management for natural voice interactions.
- Optimization & Serving: Hands-on experience with high-performance inference servers (e.g., vLLM, NVIDIA Triton, TensorRT) and optimization techniques for large-scale model deployment.
- Client to Code Translation: Demonstrated ability to act as a technical architect who can sit with stakeholders, map out domain-specific workflows (e.g., public grievance handling, CRM automation), and model them into reliable AI agents.
Bonus / Preferred Qualifications
- Deep Telephony Infrastructure: Hands-on experience with bare-metal VoIP networks, custom SIP trunks, and RTP streaming. Familiarity with managing and configuring PBX systems like Asterisk or FreeSWITCH, and handling the network latency and jitter inherent to low-level telecom systems.
Ideal Technical Stack
- Languages: Python, C++, Go (for high-performance backend components)
- AI/ML: PyTorch, vLLM, HuggingFace, LangChain/LlamaIndex, specialized ASR/TTS frameworks
- Telephony & Audio: WebRTC, standard telecom APIs (Twilio/Exotel), standard audio encoding (8kHz µ-law/A-law)
- Backend & Infrastructure: Kubernetes, Docker, gRPC, Apache Kafka, Temporal, Redis, PostgreSQL
- Observability: Prometheus, Grafana, OpenTelemetry (focusing on sub-millisecond tracing for audio/text pipelines)
📌 Lead AI Engineer – Agentic Systems & Voice AI (Hyderabad)
🏢 Yal
📍 Hyderabad