09 Sep
|
ChatBucket
|
Hyderabad
09 Sep
ChatBucket
Hyderabad
Company Description ChatBucket is a leading AI-powered multilingual communication platform that enables seamless, real-time conversations across languages. The platform eliminates the need to switch apps or rely on awkward translations, allowing natural dialogue that flows smoothly.
Designed for a world where people speak different languages, ChatBucket helps friends, businesses, customers, and global teams communicate without barriers. Using advanced AI, voice, and language technologies, it preserves context, tone, and intent in every interaction to deliver frictionless, authentic conversations.
Role Description This is a full-time, on-site Voice AI / Voice Agents Engineer role based in Hyderabad. The Voice AI / Voice Agents Engineer will design, build, and optimize voice-enabled agents that power multilingual conversations on the ChatBucket platform.
Day-to-day responsibilities include developing and integrating speech recognition, text-to-speech, and dialogue management components; collaborating with product and design teams to define voice interaction flows; and implementing features that improve latency, accuracy, and user experience. The role also involves monitoring system performance, troubleshooting issues in production, and iterating on models and pipelines based on user feedback and analytics. The engineer will work closely with cross-functional teams to ensure voice agents scale reliably and support nuanced, context-aware interactions in multiple languages.
Qualifications
Voice AI / Voice Agents Engineer
About the Role
We are looking for a Voice AI / Voice Agents Engineer to design, build, scale, and operate production-grade conversational voice systems.
You will work across speech recognition, LLMs, TTS, real-time audio processing, agent orchestration, distributed systems, and cloud infrastructure to build voice agents that can understand users, reason over context, take actions, and respond naturally with low latency.
The role requires both strong AI/ML expertise and the ability to take voice-agent systems from prototype to high-concurrency, production-scale deployments.
Responsibilities
Voice AI & Agent Engineering
- Design and develop end-to-end voice-agent architectures integrating ASR, LLMs, TTS, VAD, and agent orchestration.
- Build real-time streaming voice interactions with low latency and natural turn-taking.
- Develop agent capabilities including tool calling, function execution, memory, context management, and multi-step reasoning.
- Integrate voice agents with APIs, databases, CRMs, telephony platforms, and enterprise systems.
- Work with multilingual, noisy, and code-switching speech environments.
- Evaluate and improve ASR accuracy, TTS quality, conversational quality, interruption handling, and agent reliability.
Scaling & Distributed Systems
- Design systems capable of supporting high volumes of concurrent voice sessions.
- Build horizontally scalable services for real-time audio streaming, inference, agent orchestration, and session management.
- Design architectures for thousands to potentially millions of conversations, depending on product requirements.
- Optimize GPU and CPU utilization for large-scale model inference.
- Implement load balancing, autoscaling, connection management, queuing, caching, and backpressure mechanisms.
- Design systems that maintain predictable latency as concurrency increases.
- Optimize GPU memory, batching, continuous batching,
model serving, and inference throughput where applicable.
- Build fault-tolerant systems with graceful degradation and recovery from model, network, service, and provider failures.
- Design multi-region or highly available architectures when required.
- Optimize infrastructure and model usage for cost per conversation/minute at scale.
- Establish capacity planning and performance benchmarks for concurrent sessions.
- Identify and eliminate bottlenecks across the entire voice pipeline, from audio ingestion to model inference and response generation.
Production Engineering & Observability
- Deploy and operate AI services in production using cloud infrastructure and containerized environments.
- Build monitoring and observability for latency, throughput, concurrency, errors, model performance, GPU utilization, and infrastructure health.
- Establish SLIs/SLOs for voice-agent systems.
- Debug production issues across distributed services and real-time communication pipelines.
- Build automated evaluation and regression-testing infrastructure for voice agents.
- Develop strategies for handling provider outages, model failures, degraded network conditions, and traffic spikes.
Required Skills
- Strong Python programming and software engineering fundamentals.
- Strong understanding of machine learning, deep learning, and modern generative AI systems.
- Production experience building LLM-powered applications or AI agents.
- Practical understanding of ASR, TTS, VAD, audio processing, and streaming speech systems.
- Experience building real-time or low-latency systems.
- Robust understanding of distributed systems and backend architecture.
- Demonstrated experience scaling production services to high concurrency.
- Experience with asynchronous programming, WebSockets, REST APIs, queues, caching, and event-driven architectures.
- Experience with PyTorch or another major deep-learning framework.
- Experience with Docker, Linux, Git, and cloud infrastructure.
- Experience with production monitoring, logging, metrics, tracing, and incident debugging.
- Ability to identify bottlenecks and optimize systems across application, infrastructure, and ML inference layers.
Scaling Experience We Expect
Candidates should be able to demonstrate practical experience with questions such as:
- How would you architect a voice agent serving 10,000+ concurrent sessions?
- How would you scale streaming ASR/TTS and LLM inference independently?
- How do you prevent a sudden traffic spike from causing cascading failures?
- How would you manage thousands of persistent WebSocket/WebRTC connections?
- How do you distribute GPU inference workloads across multiple machines?
- When should you use batching, continuous batching, caching, or request queuing?
- How do you maintain low latency while increasing concurrency?
- How would you perform capacity planning for GPU inference?
- How do you design graceful degradation when an ASR, LLM, TTS, or telephony provider becomes unavailable?
- How do you measure and optimize cost per voice session?
- How would you design multi-region deployment and failover for a voice platform?
Good to Have
- Experience with WebRTC, SIP, RTP, Twilio, telephony infrastructure, or contact-center platforms.
- Experience operating high-concurrency real-time communication systems.
- Experience with Kubernetes and cloud-native infrastructure.
- Experience with AWS, GCP, or Azure.
- Experience with GPU clusters and distributed inference.
- Experience with model serving frameworks such as vLLM, NVIDIA Triton, TensorRT-LLM, or similar technologies.
- Experience with ONNX, TensorRT, quantization, and inference optimization.
- Experience with Redis, Kafka, RabbitMQ, or similar distributed messaging systems.
- Experience with PostgreSQL and distributed data systems.
- Experience designing multi-region, highly available services.
- Experience with open-source speech models such as Whisper, wav2vec 2.0, NVIDIA NeMo, or similar technologies.
- Experience with modern TTS systems and voice adaptation.
- Experience building multilingual voice systems.
- Experience with Kubernetes autoscaling and GPU scheduling.
- Experience with distributed tracing and observability platforms.
Technical Areas
AI/ML: PyTorch, Transformers, LLMs, ASR, TTS, VAD, diarization
Inference: vLLM, Triton, TensorRT-LLM, ONNX, TensorRT, quantization
Real-Time: WebSockets, WebRTC, SIP, RTP, streaming audio
Backend: Python, FastAPI, asynchronous systems, microservices
Distributed Systems: Kafka, Redis, queues, caching, load balancing, autoscaling
Infrastructure: Docker, Kubernetes, AWS/GCP/Azure, GPU clusters
Databases: PostgreSQL, Redis, vector databases
Observability: Metrics, logging, tracing, profiling, SLOs
Agent Systems: Tool calling, RAG, memory, workflow orchestration, multi-agent systems
What You Will Build
You may work on systems such as:
- High-concurrency AI phone agents
- Customer-support voice agents
- Sales and lead-qualification agents
- Appointment and scheduling agents
- Multilingual conversational assistants
- Enterprise voice copilots
- AI-powered contact-center platforms
- Accessibility-focused voice systems
- Large-scale autonomous voice-agent platforms
Ideal Candidate
We are looking for an engineer who can bridge AI research, production engineering, and distributed systems.
You should be able to take a voice-agent prototype and answer:
> How do we make this reliable, low-latency, cost-efficient, and capable of serving thousands of concurrent users?
The ideal candidate understands not only how to build an individual AI agent, but also how to deploy, observe, optimize, and scale the entire system in production.
Education & Experience
* Bachelor's or Master's degree in Computer Science, AI, ML, Speech Processing, or a related field is preferred.
- 3+ years of relevant engineering experience preferred.
- Strong candidates with significant production projects, open-source contributions, or research experience are encouraged to apply.
- Experience operating large-scale AI or real-time systems in production is highly valued.
Success in This Role
Success means building voice AI infrastructure that is:
- Low latency
- Highly concurrent
- Horizontally scalable
- Fault tolerant
- Observable
- Cost efficient
- Production reliable
You will help transform voice agents from individual AI prototypes into large-scale, production-grade conversational systems.
📌 Voice AI / Voice Agents Engineer (Hyderabad)
🏢 ChatBucket
📍 Hyderabad