Technical Program Manager, Speech Data (Bengaluru)

Technical Program Manager, Speech Data (Bengaluru)

27 Sep
|
Wiingy AI
|
Bengaluru

27 Sep

Wiingy AI

Bengaluru

About the role

You will plan and run our speech data projects from start to finish. You will design each corpus, write its specifications, set up the collection and annotation pipeline, and check data quality using scripts and automated checks. You will also find and manage the people who produce the data, including speakers, annotators, linguists, and vendors, and you will run collection so that each dataset is delivered on time, on budget, and with proper consent.

You should understand how speech models are trained, so that the data you design and deliver works well in a buyer's training pipeline.

You do not need to be an expert in every language or country. You need solid experience in a few languages and the ability to apply what you know to new ones, working with native-speaker linguists and local partners.

Work Policy: Onsite

Location: Bangalore

What you will do

Corpus design and program planning

- Design each corpus based on how the data will be used in training. This includes deciding what speakers should say (scripts, prompts, topics, or conversation scenarios), the speaker mix by language, dialect, age, gender, and region, and the recording conditions, such as studio, home, phone, or noisy environments.
- For text-to-speech data, make sure scripts give good coverage of the sounds and sentence types in each language.
- Turn corpus designs into clear specifications, including audio quality, file formats, metadata, and delivery format.
- Plan each project, including timeline, budget, staffing, and cost per usable hour of audio.
- Identify risks early, such as low speaker supply for a dialect or weak tools for a language, and plan around them.

Collection pipeline and tooling

- Set up the collection and annotation workflow, including recording apps, annotation tools, storage, and review steps.
- Find and recruit speakers who match each corpus design, directly or through vendors and local partners.
- Run contributor onboarding, payments, and support, and prevent fraud such as the same person recording under several accounts.




- Make sure every contributor gives proper consent, and that each recording can be traced to its contributor and consent record.
- Recruit, train, and manage annotators and linguists, and work with linguists to write transcription guidelines for each language.
- Use speech models to speed up work, for example by pre-transcribing audio for human correction.

Quality and measurement

- Run automated quality checks on audio and transcripts, and review samples by hand.
- Analyse data with scripts to track quality, speaker coverage, rejection rates, and progress against the plan.
- Run small pilot batches before full collection, and revise the corpus design, specifications, or guidelines based on the results.
- Prepare training, validation, and test splits that avoid problems such as the same speaker appearing in both training and test data.
- Where useful, run simple training or fine-tuning experiments to check that the data improves model performance.

Delivery and coordination

- Prepare datasets for delivery, including file formats, metadata, and documentation that describes how the data was collected.
- Coordinate speakers, annotators, linguists, vendors, and engineers so that each project stays on schedule and on budget.
- Join technical conversations with AI labs alongside our founders, and handle their feedback on delivered data.

What you must have
- At least 1 year working on speech or audio data. This can be data collection, annotation, transcription, or speech model development.
- Experience designing a speech corpus or dataset, including deciding what content to record and which speakers to include.




- Explicit understanding of how ASR and TTS models are trained, including how data is cleaned, segmented, aligned, split, and used in training and evaluation.
- Experience running at least one data collection or annotation project from planning to delivery.
- Experience with speech data in at least two languages.
- Experience recruiting and managing contributors, annotators, or field workers at scale, and managing external vendors.
- Experience writing or applying transcription or annotation guidelines.
- Understanding of how audio quality is measured and why it matters for model training.
- Basic Python or similar scripting skills, enough to analyse data, run simple audio checks, and work with engineers.

What will help your application
- Experience at a speech AI company, a speech research lab, or a data vendor that sells speech data.
- Hands-on experience training or fine-tuning a speech model.
- Experience with Indian languages, or with languages in other regions such as Southeast Asia, the Middle East, Africa, Europe, or Latin America.
- Experience with text-to-speech data or conversational speech.
- Experience with annotation tools such as Label Studio.
- Knowledge of consent and privacy rules for voice data.

About Wiingy AI

Wiingy AI provides human data and AI evaluation services to AI labs and AI companies. We supply expert human judgment where automated methods are not good enough, including evaluating model outputs, producing feedback data for model training, and building curated datasets.

Wiingy AI is part of Wiingy, the established tutoring platform whose network of 5,000+ credentialed, vetted subject-matter experts musicians with formal training, language and STEM educators, domain specialists already delivers expert instruction at scale. Wiingy AI turns that network into the judgment layer of AI training.

This role sits with Wiingy AI, a distinct business from the Wiingy tutoring platform, powered by the same expert network.

📌 Technical Program Manager, Speech Data (Bengaluru)
🏢 Wiingy AI
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: technical program manager, speech data (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: technical program manager, speech data (bengaluru) / bengaluru