About role:
You'll be the engineering owner on the speech team, working alongside two ML engineers and building everything around the models — ingestion, training infrastructure, real-time serving, and the integration into our existing platform.
This is a backend engineering role. You don't need to train models. You need to build the systems that make trained models practical in production.
Responsibilities:
- Build the audio data pipeline: mine our call archive, transcode, resample, segment, deduplicate and quality-filter at scale
- Build and operate real-time inference services with hard latency targets, including streaming, cancellation and mid-utterance interruption
- Integrate speech services into our existing Java-based platform and telephony infrastructure
- Instrument the full latency budget end to end and find where the milliseconds go
- Stand up training infrastructure — GPU scheduling, checkpointing, experiment tracking, reproducibility
- Own compute cost and concurrency economics: how many simultaneous calls per GPU, and how to improve it
- Support on-premise deployment for clients who can't send data outside their network
RequirementsMust haves:
- 3–4 years building and operating production backend systems
- Strong Java — you've owned services in production, not just contributed to them
- Working Python — enough to build data pipelines and integrate with ML tooling
- Real-time or low-latency systems experience: streaming APIs, WebSocket or gRPC streaming, concurrency, backpressure
- Data pipelines at scale (Airflow, Dagster, Spark or equivalent)
- Docker and Kubernetes in production
- Cloud infrastructure (AWS/GCP/Azure)
- Comfortable debugging performance: profiling, latency percentiles, throughput under load
Nice to have:
- Serving ML models in production (Triton, vLLM, TorchServe)
- Audio tooling — ffmpeg, sox, codecs, resampling
- Telephony — SIP, Asterisk/FreeSWITCH, media servers, narrowband codecs
- MLOps tooling: MLflow, Weights & Biases, DVC
- GPU-aware infrastructure work
📌 Senior Backend Engineer - Speech Platform (Noida)
🏢 Rezo.AI
📍 Noida