16 Sep
|
OnebyZero Pte
|
Bengaluru
16 Sep
OnebyZero Pte
Bengaluru
Core Product – Voice & Real-Time Infrastructure
Location: Bangalore, India / Hybrid
Employment Type: Full-time
Team: Core Product Engineering
About OneByZero
OneByZero builds AI-native platforms that help enterprises in regulated industries — banking, telecommunications, and retail — move beyond AI experimentation into production-ready, governed AI workforces. Our AWS-native platform deploys “governed AI coworkers”: autonomous agents that handle real business processes while maintaining full auditability, data sovereignty, and human oversight, typically moving from pilot to production in weeks rather than quarters.
We are not building a straightforward wrapper around a model — we are building a commercial platform that negotiates, handles complex workflows, and converses naturally with humans at enterprise scale.
The Mission
At our core, we are solving for seamless human-machine interaction — using real-time voice as the ultimate interface. We are building an open-source-driven, AWS-native AI platform designed to deploy as an agentic digital coworker for enterprises.
Your mission is to own the zero-to-one build-out of the ultra-low-latency audio architecture that makes this interaction feel instantaneous, bridging the gap between modern LLM orchestration and legacy enterprise telephony.
What You Will Do
Architect the core pipeline
- Design and maintain a highly concurrent, bi-directional audio streaming infrastructure using WebRTC, WebSockets, and gRPC.
- Handle the complexities of network traversal including STUN/TURN and ICE candidate negotiation.
- Manage packet loss and transcoding across relevant codecs such as Opus and G.711 μ-law/A-law.
Bridge the legacy gap
- Build signaling and CTI (Computer Telephony Integration) connectors required to interface seamlessly with legacy on-prem and cloud CCaaS platforms such as Avaya AES, Genesys, and Cisco.
- Manage call control signals to agent desktops alongside real-time RTP media extraction.
Manage state and concurrency at scale
- Architect a distributed, cloud-native system on AWS using technologies such as EKS and ElastiCache/Redis.
- Maintain stateful streaming sessions and handle thousands of concurrent WebSocket/WebRTC connections without race conditions or dropped packets.
Drive the product rollout
- Transition the voice pipeline from its current build into a hardened, production-ready system.
- Support high-volume, multi-tenant enterprise deployments using modern CI/CD and Infrastructure as Code (Terraform).
Kill latency
- Optimize every millisecond of the audio transport layer, STT, and TTS handoffs.
- Keep system latency well under 500ms at scale.
Solve the “human” problems
- Build robust Voice Activity Detection (VAD) and endpointing.
- Handle natural pauses, background noise, and real-time human interruptions (barge-in).
Integrate with the “brain”
- Work closely with product engineers to ensure audio streams and CTI signals feed cleanly into our LLM state management and tool-calling infrastructure.
Requirements
What You Must Have
- Deep, hands-on experience building low-latency, real-time media servers or audio pipelines at scale, heavily utilizing open-source frameworks such as LiveKit, Janus, Mediasoup, or GStreamer.
- Strong systems programming proficiency in Go, Rust, C++, or Python for handling heavy concurrency and network I/O.
- Knowledge of AWS networking and compute primitives, with a track record of deploying stateful, scalable streaming architectures natively in the cloud.
- Proven experience working with enterprise telephony, including SIP trunking, RTP streams, and CTI protocols.
- Expertise in managing ICE connection states, stream buffering, jitter buffers, and signal processing basics.
- 5+ years of experience in backend, infrastructure, or media/telephony engineering, with a portion of that time spent owning a system end-to-end from design through production.
Nice to Have
- Prior experience integrating speech-to-text (STT) and text-to-speech (TTS) providers into a real-time conversational pipeline.
- Familiarity with LLM orchestration frameworks and function/tool-calling patterns.
- Experience operating Kubernetes (EKS) in production, including autoscaling stateful workloads and zero-downtime deployments.
- Exposure to observability tooling for real-time systems, including distributed tracing, RTP/media quality metrics, and latency dashboards.
- Background in contact center technology, IVR systems, or telecom carrier integrations.
- Experience contributing to or maintaining open-source real-time communication projects.
What Success Looks Like
- A production-hardened voice pipeline running at sub-500ms end-to-end latency, supporting thousands of concurrent sessions across multiple enterprise tenants.
- CTI connectors live with at least one major CCaaS platform — Avaya, Genesys, or Cisco — enabling seamless agent-desktop handoff.
- A distributed session-state architecture on AWS that survives node failures and scales horizontally without dropped calls.
- Natural, low-friction conversations with accurate barge-in handling and endpointing that feels human, not robotic.
Benefits
Why Join OneByZero
- Own a zero-to-one build: This is a foundational, high-ownership role shaping the core real-time infrastructure of the product.
- Work at the intersection of telecom, distributed systems, and applied AI: Solve problems most engineers never get exposed to.
- Ship into real enterprise deployments in regulated industries: Gain direct visibility into production impact.
📌 Senior Voice Pipeline Engineer (Bengaluru)
🏢 OnebyZero Pte
📍 Bengaluru