17 Sep
|
TalixoHR
|
Bangalore Metropolitan Area
17 Sep
TalixoHR
Bangalore Metropolitan Area
Experience: 3–4 years
Employment: Full-time
Location: HSR Layout Bengaluru | In office
About the Role
We are looking for a Senior Voice AI / Speech ML Engineer to join our AI team and build the next generation of speech and voice intelligence systems.
This role is for an engineer/researcher who has hands-on experience training large-scale speech or language foundation models , not just integrating or fine-tuning existing models.
You will work across ASR, TTS, speaker identification, VAD and conversational AI , taking models from large-scale data preparation and training through evaluation, optimization and production deployment.
If you have personally trained/pre-trained a foundation model and have strong ASR/TTS experience, we would like to hear from you.
What You'll Work On
- Design, train, pre-train and fine-tune foundation models for Voice AI.
- Develop and train ASR and/or TTS models for real-world speech applications.
- Work with large-scale speech, audio and text datasets .
- Build data preparation, preprocessing, augmentation and model-training pipelines.
- Design experiments and evaluate models across accuracy, robustness and generalization.
- Improve WER/CER, speech quality, latency and inference efficiency .
- Work with Transformer-based architectures and modern speech-model architectures.
- Train models using GPUs and distributed computing infrastructure .
- Debug training instability, data issues and model-performance bottlenecks.
- Optimize trained models for production, including quantization and inference optimization .
- Collaborate with research and engineering teams to move models from experimentation to production.
Must-Have Experience
- 3–4 years of hands-on experience in Speech ML, Speech AI, NLP, Deep Learning or Machine Learning.
- Actual foundation-model training/pre-training experience.
- Hands-on experience developing or training ASR and/or TTS models .
- Strong understanding of Transformers and deep learning .
- Strong Python programming skills.
- Hands-on experience with PyTorch and/or TensorFlow .
- Experience with GPU-based model training and large datasets.
- Strong understanding of model evaluation, experimentation and optimization.
Strongly Preferred Experience with one or more of:
- Whisper
- wav2vec 2.0
- HuBERT
- Conformer / FastConformer
- NVIDIA NeMo
- SpeechBrain
- Hugging Face
- F5-TTS / VITS / similar TTS architectures
- Speech or audio foundation models
- Multilingual / low-resource speech
- Indian-language speech datasets
- Distributed training using DeepSpeed, FSDP, Ray Train or similar
- Audio preprocessing, augmentation, annotation and dataset-quality pipelines
- Production deployment of speech models
What We Are Specifically Looking For The ideal candidate has worked on the model itself , rather than only building applications around existing models.
Strong fit:
Foundation-model pretraining + ASR/TTS + Transformers + PyTorch + GPU/distributed training + large-scale speech data
Not the right fit:
- Generic ML Engineer without speech experience
- GenAI / RAG Engineer
- Prompt Engineer
- LLM application developer
- Candidates who only consume OpenAI/Claude/Gemini APIs
- Candidates who only fine-tune existing pretrained models
- Conversational AI developers who integrate speech APIs/Riva/voice APIs but don't train speech models
- Candidates with only NLP/LLM experience and no ASR/TTS
Why Join
- Work on core Voice AI / Speech ML technology , not only application-layer GenAI.
- Solve challenging problems involving large-scale model training, speech data and inference optimization .
- Work closely with AI/ML engineers and researchers on production-grade models.
- Opportunity to work on multilingual and real-world speech applications in a high-growth AI environment.
📌 Senior Voice AI / Speech ML Engineer (Bangalore Metropolitan Area)
🏢 TalixoHR
📍 Bangalore Metropolitan Area