- Fine-tune and adapt open-source speech models on our proprietary call audio
- Build training and evaluation pipelines for multilingual and code-mixed speech
- Own model quality against metrics that matter operationally — entity and numeric accuracy, latency to first response — not just aggregate error rates
- Design and run data curation at scale: pseudo-labelling, speech enhancement, quality filtering on messy real-world audio
- Work with our linguist on text normalisation and pronunciation handling
- Evaluate candidate architectures, make the call with evidence, and ship the result to production with the platform team
Requirements
Must have skills:
- 3–4 years in ML, with at least 18 months on speech or audio specifically
- Robust Python and PyTorch; comfortable reading a paper and implementing it
- Hands-on experience fine-tuning at least one production speech model
- Solid grasp of speech fundamentals — mel-spectrograms, acoustic models and vocoders, encoder-decoder vs transducer architectures, evaluation methodology, sampling rates and what they cost you
- Understanding of how modern speech systems are actually built: self-supervised encoders, neural audio codecs, LM-based generation, flow matching
- Experience with genuinely messy audio, not only clean benchmark datasets
Nice to have:
- NeMo, ESPnet, SpeechBrain, or Coqui
- Telephony-band or contact-centre audio
- Multilingual or code-switched speech work
- LoRA/PEFT, distributed training
- Open-source contributions or publications in speech
📌 ML Engineer - Speech (Noida)
🏢 Rezo.AI
📍 Noida
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.