• Conduct objective and subjective speech evaluation.
• Work closely with speech-data specialists.
• Work with ML systems engineers on inference optimisation.
• Document experiments and model-performance results.
REQUIRED SKILLS
• Python
• PyTorch
• Deep Learning
• Audio ML
• Speech processing
• TTS training and fine-tuning
• Audio preprocessing
• PCM/WAV fundamentals
• Sampling rates
• Spectrograms
• GPU training
• Mixed precision
• Dataset preparation
• Model evaluation
• Linux
• Git
SIGNIFICANT CONCEPTS
Candidates should understand
• Prosody
• Phonemes
• Speaker embeddings
• Vocoders
• Streaming synthesis
• Real-time factor
• Time to first audio
• Speaker conditioning
EXPERIENCE WITH ANY OF THE FOLLOWING IS VALUABLE
• XTTS
• VITS
• FastSpeech
• StyleTTS
• Parler-TTS
• Neural codecs
• Diffusion-based speech models
• Flow-based TTS
• LoRA
• ONNX
• TensorRT
EDUCATIONAL QUALIFICATION
Preferred
B.Tech/B.E./M.Tech/M.Sc. in Computer Science, Electronics, AI/ML, Signal Processing, Speech Technology, Computational Linguistics or a related field.
Research candidates with lower corporate experience may be considered if they can demonstrate meaningful TTS model-development experience.
RESEARCH EXPERIENCE IS WELCOME
We encourage applications from candidates who have:
- Published relevant research
- Released models on Hugging Face
- Contributed to open-source speech projects
- Worked in speech research laboratories
- Independently trained or fine-tuned speech models
API integration with commercial TTS providers alone will not be considered sufficient model-development experience.
APPLICATION DETAILS
Please include
- Updated CV
- Current location
- Current CTC
- Expected CTC
- Notice period
- GitHub / Hugging Face links
- Publications, if applicable
- Audio samples or model demonstrations, where available
- Short description of relevant speech-model work
CONFIDENTIALITY
This position involves confidential research and product development. Specific product details, datasets and internal architecture will be disclosed only at the appropriate stage of the hiring process and may require an NDA.
📌 Machine Learning Engineer - TTS & Expressive Speech (Kolkata Metropolitan Area)
🏢 Ethical Den
📍 Kolkata Metropolitan Area
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.