04 Sep
|
Wiingy
|
Bengaluru
Applied Research Lead - Voice AI & Benchmarking
Location: Bengaluru, India
Employment: Full-time
Reporting to: Co-founder
About Wiingy AI
Wiingy AI is an expert-led human data and model evaluation company. We help teams building AI systems create high-quality datasets, conduct rigorous human evaluations and understand where their models fail.
Our voice work spans text-to-speech, automatic speech recognition, multilingual speech, accents, dialects, code-switching, pronunciation, prosody and real-world acoustic conditions.
We are building a technical research function that connects speech-model science with large-scale human data operations. A central part of this function will be an independent global benchmark that measures how voice models perform across languages, accents, use cases and real-world conditions.
The role
We are looking for an Applied Research Lead, Voice AI & Benchmarking to become Wiingy AI’s internal authority on speech models, speech datasets and model evaluation.
You should understand how modern text-to-speech, automatic speech recognition and speech-to-speech systems are built, trained, fine-tuned and evaluated.
Your primary mandate will be to build, launch and continuously advance Wiingy AI’s global voice benchmark. The benchmark should become a trusted, transparent and technically rigorous way to compare voice models across languages, accents, tasks and real-world conditions.
You will own the benchmark as a continuing research program rather than a one-time report. This includes its methodology, evaluation datasets, human evaluation system, technical harness, release cadence, public leaderboard and research roadmap.
Primary mandate: Build a global voice benchmarkYou will:
- Define what the benchmark measures, which users it serves and how it remains differentiated, credible and useful as voice technology changes.
- Establish the initial scope across TTS and ASR, with a roadmap covering speech-to-speech systems and real-time voice agents.
- Design representative tasks, prompts, test cases, domains and adversarial conditions.
- Create meaningful global coverage across languages, countries, accents, dialects, code-switching, speaker characteristics, speaking styles and recording environments.
- Define how models, model versions, voice IDs, inference parameters and system configurations are selected and disclosed so comparisons remain fair and reproducible.
- Combine objective metrics with blind human evaluation, using dimensions, scales, pairwise comparisons, error taxonomies and aggregation methods appropriate to each task.
- Determine when evaluations require speech experts, native speakers, domain experts or general listeners.
- Build rater qualification, calibration, gold-task, monitoring, disagreement-resolution and adjudication procedures.
- Establish standards for sample size, randomization, counterbalancing, confidence intervals, significance testing,
inter-rater reliability and repeatability.
- Build or supervise the technical evaluation harness used to run models, preserve inference settings, randomize outputs, collect judgments, calculate metrics and generate reproducible reports.
- Publish clear methodology, dataset documentation, limitations, results and a versioned leaderboard that external researchers and model builders can scrutinize.
- Continuously advance the benchmark by tracking model releases, adding languages and failure categories, refreshing test sets, improving metrics and publishing regular updates without breaking historical comparability.
- Protect the benchmark from contamination, overfitting, selective reporting, commercial influence and undisclosed methodology changes.
- Use benchmark findings to identify systematic model failures and produce original research, technical reports and failure analyses.
What you will ownBeyond the benchmark, you will:
- Serve as Wiingy AI’s go-to person for TTS, ASR and speech-data questions across research, operations, product and go-to-market teams.
- Maintain a working understanding of current speech-model architectures, training and fine-tuning methods, inference controls, failure modes and evaluation practices.
- Design collection and annotation specifications for pre-training, fine-tuning, supervised evaluation and preference data.
- Define standards for recording environments, speaker selection, audio formats, segmentation, transcription, text normalization, pronunciation lexicons, phonetic labels, alignment, metadata and quality assurance.
- Design TTS evaluations covering naturalness, intelligibility, pronunciation, prosody, speaker similarity, emotion, style adherence, controllability, artifacts, robustness and latency.
- Design ASR evaluations using WER, CER and task-specific error taxonomies, including named entities, numbers, accents, dialects, code-switching, noisy speech, far-field speech and domain vocabulary.
- Build or supervise lightweight pipelines for model inference, audio processing, dataset inspection, metric computation, error analysis and benchmark reporting.
- Translate customer objectives into rigorous evaluation plans and dataset specifications.
- Incorporate speaker consent, licensing, privacy, provenance, bias, safety and appropriate dataset use into every research design.
What you should bring
- At least three years of relevant experience in speech AI, machine learning, computational linguistics, audio ML or a closely related field.
- Hands-on experience training, fine-tuning or deeply evaluating at least one TTS or ASR system.
- Strong understanding of speech pipelines, including audio preprocessing, feature representations, tokenization or phonemes, alignment, decoding or generation and inference-time controls.
- Practical experience designing speech datasets, evaluation sets or model benchmarks.
- Working knowledge of TTS and ASR evaluation metrics and their limitations, including human evaluation methodology.
- Robust experimental-design and statistics fundamentals, including sampling, bias, variance, confidence intervals, significance testing, agreement measurement and reproducibility.
- Ability to work in Python and with common speech or ML tooling such as PyTorch, Hugging Face, Whisper-family tools and audio-processing libraries.
- Ability to explain technical decisions clearly to researchers, operators, product teams, customers and non-technical stakeholders.
- A master’s degree or PhD in a relevant discipline is preferred. Equivalent industry research or engineering experience is equally valuable.
Especially valuable
- Experience designing, launching or maintaining a public benchmark, shared task, model leaderboard or reproducible evaluation harness.
- Experience with multilingual or low-resource speech, accents, dialects or code-switching.
- Experience evaluating multiple commercial or open-source models under controlled and comparable inference settings.
- Knowledge of forced alignment, grapheme-to-phoneme systems, IPA, pronunciation modelling, text normalization or speech-quality assessment.
- Experience with preference data, RLHF or other post-training methods for speech or generative-audio models.
- Experience working with large annotation operations, external data vendors or distributed expert workforces.
- Published research, open-source contributions or strong technical writing in speech or audio ML.
ABOUT WIINGY AI Wiingy AI is a new business built on Wiingy, the established tutoring platform whose network of 5,000+ credentialed, vetted subject-matter experts musicians with formal training, language and STEM educators, domain specialists already delivers expert instruction at scale. Wiingy AI turns that network into the judgment layer of AI training.
The frontier is bottlenecked on judgment, not compute. Labs and applied-AI teams can crowdsource clicks; they cannot crowdsource expertise and that is exactly what we supply: expert human evaluation, preference/RLHF data, benchmark construction, and pedagogical QA. We compete on proven expertise and rigor: published inter-annotator agreement, calibrated rubrics, defensible methodology in a market where the premium judgment layer is still being defined.
This role sits with Wiingy AI, a distinct business from the Wiingy tutoring platform, powered by the same expert network.
📌 Applied Research Lead — Voice AI & Evaluations (Bengaluru)
🏢 Wiingy
📍 Bengaluru