16 Aug
|
Alion
|
Anupgarh
You'll own the distributed systems that orchestrate parallel voice calls at massive scale, manage campaigns across millions of accounts, and keep our real-time voice AI pipeline humming under pressure. This isn't glue-code work; it's designing globally distributed systems where every millisecond and every dropped connection counts, requiring you to achieve massive horizontal scaling while strictly maintaining end-to-end low latency.
Responsibilities
- Voice AI at Scale: Design and operate the infrastructure that launches, monitors, and recovers thousands of concurrent voice sessions. You will manage parallel call orchestration, connection pooling, and real-time health tracking.
- Campaign Management Engine: Build systems that allow teams to configure, schedule, and throttle outbound campaigns across millions of accounts with complex retry logic and DNC (Do Not Call) compliance.
- Real-Time Pipeline Reliability: Own the availability and latency of the ASR LLM TTS pipeline. You will instrument, trace, and optimise the "hot path" to ensure calls stay within strict latency SLAs.
- Platform and Data Architecture: Evolve our backend services, queues, and data stores to handle 10x growth. Lead design reviews and make critical build-vs-buy decisions for core infrastructure.
- Integrations and APIs: Design robust APIs and webhook systems that connect our platform to client CRMs, payment gateways, and telephony providers with high throughput and reliable backoff semantics.
- Operational Excellence: Champion observability, incident response, and capacity planning. Build the dashboards and runbooks that ensure system stability during peak loads.
Requirements
- Distributed Systems: Deep hands-on experience building and operating distributed backends, focusing on fault tolerance, consensus, partitioning, and high-availability failure modes. You should be adept at designing for consistency models (e. g., CAP theorem trade-offs), managing stateful services, and optimising for latency-throughput trade-offs in high-volume production environments.
- Languages: Proficiency in Python, Go, or Java/Kotlin. You should have robust opinions on concurrency primitives and async I/O patterns.
- High-Throughput Infra: A proven track record with Apache Kafka, RabbitMQ, or SQS.
Experience in stream processing, real-time pipelines, and event-driven architectures.
- Data Stores: Production experience with PostgreSQL, Redis, DynamoDB, or similar. You know when to reach for SQL vs. NoSQL.
- Cloud and Orchestration: Fluent in AWS or GCP. Comfortable with Kubernetes, Terraform, and CI/CD pipelines.
- Observability: Ability to build systems that are "debuggable by default" using distributed tracing and SLO-driven alerting (Datadog, Grafana, Prometheus).
- API Design: Experience scaling RESTful and/or gRPC APIs serving high-QPS traffic with proper versioning and auth.
- Ownership Mindset: 5+ years of building production systems. You are comfortable making architectural calls, leading RFCs, and mentoring junior engineers.
Bonus Points
- Experience with Telephony / VoIP / SIP or real-time media streaming.
- Knowledge of LLM serving infrastructure or ASR/TTS pipelines.
- Experience in Fintech / Collections domains or TCPA/RBI compliance.
- Familiarity with WebSockets or WebRTC.
📌 Senior Backend Engineer (Anupgarh)
🏢 Alion
📍 Anupgarh