ML Engineer - Speech
Responsibilities
- Fine-tune and adapt open-source speech models on our proprietary call audio
- Build training and evaluation pipelines for multilingual and code-mixed speech
- Own model quality against metrics that matter operationally - entity and numeric accuracy, latency to first response - not just aggregate error rates
- Design and run data curation at scale: pseudo-labelling, speech enhancement, quality filtering on messy real-world audio
- Work with our linguist on text normalisation and pronunciation handling
- Evaluate candidate architectures, make the call with evidence, and ship the result to production with the platform team
Requirements
- 3-4 years in ML, with at least 18 months on speech or audio specifically
- Strong Python and PyTorch; comfortable reading a paper and implementing it
- Hands-on experience fine-tuning at least one production speech model
- Solid grasp of speech fundamentals - mel-spectrograms, acoustic models and vocoders, encoder-decoder vs transducer architectures, evaluation methodology, sampling rates and what they cost you
- Understanding of how up-to-date speech systems are actually built: self-supervised encoders, neural audio codecs, LM-based generation, flow matching
- Experience with genuinely messy audio, not only clean benchmark datasets
- Multilingual or code-switched speech work
- LoRA/PEFT, distributed training
- Nice to have: NeMo, ESPnet, SpeechBrain, or Coqui
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 ML Engineer - Speech (Noida)
🏢 Rezo
📍 Noida