ML Research Engineer - Speech/Audio (India)

ML Research Engineer - Speech/Audio (India)

02 Oct
|
Eliteeye Consulting
|
India

02 Oct

Eliteeye Consulting

India

About the Role:

You will own a research direction (STT, TTS or speech-to-speech) end to end. That means designing new architectures, pre-training foundation models from scratch on large multilingual audio, and getting them into production, while mentoring engineers on the team.

What you will do:

- Lead one research track and set its roadmap with the CTO and product leads, tied to customer metrics (accuracy, naturalness, latency, cost per minute).
- Design novel architectures and pre-train speech foundation models from scratch on large multilingual corpora (100k+ hours), including self-supervised pre-training, tokenizer and codec design, and scaling experiments. Fine-tuning and continued training are tools, not the whole job.
- STT: streaming ASR with low time-to-first-token, code-switch robustness, contextual biasing for names and entities, domain adaptation for BFSI and healthcare.
- TTS: expressive, low-latency streaming TTS for Indian languages; voice cloning, prosody and emotion control, pronunciation of numbers, dates and mixed-script text.
- Speech-to-speech: Full-duplex conversational models built in-house: our own audio tokenizers and codecs, speech-LLM pre-training and alignment, interruption and turn-taking, end-of-utterance prediction.
- Build the data strategy: sourcing, licensing, synthetic data generation, labelling quality and consent-compliant use of production audio.
- Define evaluation standards and benchmarks that predict real call outcomes, and run honest comparisons against external vendors.
- Make serving trade-offs with the platform team: distillation, quantisation, GPU concurrency per model, on-prem footprints.
- Mentor 1 - 3 engineers; review experiment designs and code.




- Publish or open-source selectively where it strengthens our position.

What you bring:

- 3 - 6 years in ML, with at least 2 - 3 years focused on speech (ASR, TTS, speaker/voice or audio-language models).
- MS or PhD in CS, EE or related field, or equivalent depth shown through shipped work.
- Track record of pre-training at least one speech or audio model from scratch (not only fine-tuning) that reached production or strong published results.
- Deep knowledge of contemporary architectures: Conformer/Zipformer, RNN-T/TDT, encoder-decoder ASR, flow-matching and diffusion TTS, neural codecs (EnCodec, DAC, Mimi), speech-LLMs.
- Large-scale pre-training experience: distributed training (FSDP/DeepSpeed/Megatron) on multi-node GPU clusters, 100k+ audio hours, scaling laws and compute budgeting, data loading at scale, debugging loss spikes and unstable runs.
- Strong experimental design: ablations, significance, avoiding test-set leakage.
- Clear written and spoken communication with engineers, product and customers.

Nice to have:

- Publications at Interspeech, ICASSP, ACL/EMNLP, NeurIPS or similar.
- Experience with Indic languages, low-resource languages or code-switching.
- Built streaming or real-time speech systems with strict latency SLAs.
- Experience with RL or preference tuning for speech (e.g. naturalness rewards) or LLM post-training.
- Contributions to open-source speech toolkits.

What success looks like in 12 months :

- Your track has shipped a model pre-trained from scratch by Blue Machines that beats the vendor it replaces on our benchmark and reduces cost or latency.
- A clear, reproducible training and evaluation stack for your area.
- Engineers you mentor are running experiments independently.

📌 ML Research Engineer - Speech/Audio (India)
🏢 Eliteeye Consulting
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: ml research engineer - speech/audio (india) / india