Data Engineer (Speech/Language Data) (Bengaluru)

Data Engineer (Speech/Language Data) (Bengaluru)

11 Sep
|
Gnani Innovations
|
Bengaluru

11 Sep

Gnani Innovations

Bengaluru

Responsibilities

Data Engineer / Analyst - Speech & Language Data

Location: Bengaluru

Type: Full-time

Function: AI Team

Role overview

Every speech model we ship is downstream of a decision someone made about data. Which hours got recorded. Which transcripts were trusted. Whether the Tamil test set accidentally shares speakers with the Tamil training set. Whether a number was written as or 25. Those decisions set the ceiling on what our ASR and TTS models can ever achieve - and no amount of GPU time recovers a corpus that was built carelessly. Training accurate ASR across 22+ Indian languages takes tens of thousands of hours of audio. Training TTS voices that sound human takes carefully recorded, tightly aligned, phonetically balanced speech. Right now that data lives across many places and many heads. We want one person to own it.

This role is the data owner for Gnani s AI team. You source it, clean it, label it, version it, catalogue it, and defend its provenance - and you are the single point of contact when any researcher needs hours, a benchmark, or an answer about where something came from. It is a hands-on engineering and analysis role with real judgment attached: you will listen to audio, read transcripts, and be the person who notices that something is wrong before a model trains on it.

Core mandate

Sourcing & acquisition Find, evaluate, and bring in speech and text data across 22+ Indian languages - public, licensed, customer, and commissioned

Pipelines & data quality

Build the cleaning, segmentation, alignment, and validation pipelines that turn raw audio into trainable hours

Annotation & benchmarks

Own transcription and labelling quality, and maintain the frozen evaluation sets the whole team is judged on

Data ownership & governance

Be the single point of contact for data - catalogued, versioned, consented, and compliant by default

What you'll drive

Sourcing & acquisition

- Be permanently on the hunt for data - track new public Indic speech and text corpora as they release, evaluate them for quality and licence terms, and bring the good ones in quick
- Know the landscape cold: the AI4Bharat ecosystem (IndicVoices, Kathbath, Shrutilipi, GramVaani, IndicTTS), LIMMITS, SYSPIN, Rasa, Common Voice, FLEURS, and whatever lands next - and maintain an honest view of what each is actually good for
- Assess and onboard customer-contributed data - what consent covers it, what the DPA permits, what has to be de-identified before it can be used, and what cannot be used at all




- Spec and manage commissioned data collection and studio recording - script design for phonetic and dialect coverage, speaker recruitment criteria, session QC, and vendor throughput
- Keep a running gap analysis: hours per language, per dialect, per channel, per domain - so the team argues about priorities with numbers instead of impressions

Pipelines, cleaning & curation

- Build and own the pipelines that turn raw audio into trainable hours - format conversion and resampling, loudness normalisation, VAD-based segmentation, diarisation, and forced alignment
- Automate quality screening at scale: SNR estimation, clipping and DC offset, truncated utterances, silence-heavy segments, language-ID mismatches, and audio-transcript drift
- De-duplicate seriously - audio fingerprinting and near-duplicate text detection, so the same recording is not both trained on and evaluated on
- Own the text side, which for Indic data is most of the work: orthographic convention, numeral form, transliteration and romanised code-switch, punctuation, casing, inverse text normalisation, and lexicon maintenance - inconsistent conventions show up as WER that has nothing to do with the model
- Run large-scale pseudo-labelling and weak supervision - transcribe unlabelled audio with existing models, filter on confidence and cross-model agreement, and route the uncertain cases to human review
- Version and catalogue everything: reproducible dataset snapshots, training manifests, dataset cards recording source, licence, consent basis, and known limitations

Annotation, listening & quality

- Own transcription and labelling quality end to end - write the annotation guidelines, train annotators against them, and revise them when reality disagrees
- Listen to audio and annotate yourself, as needed. Sampling real data is how you find the problems dashboards hide, and it is part of this job rather than beneath it
- Measure annotation quality properly - gold sets, inter-annotator agreement, WER between independent transcriptions - and manage vendors on quality and cost per hour, not volume alone
- Build the TTS-specific quality gates: alignment tightness,



transcript fidelity, speaker and style consistency, prosody and emotion tagging, and codec round-trip checks
- Coordinate listening panels for subjective evaluation (MOS / CMOS / preference tests) - recruitment, screening, sample randomisation, and result analysis

Benchmarks & evaluation sets

- Own Gnani s internal benchmark suite for ASR and TTS - curated, frozen, versioned, speaker-disjoint from training data, and representative across language, dialect, accent, channel, and domain
- Guarantee split hygiene. No test speaker in train, no leaked utterance, no silently mutated eval set - and be willing to block a result that violates it
- Track external and public benchmarks, reproduce them faithfully, and flag when a published comparison is not apples to apples
- Analyse results, not just produce them - slice error by language, speaker, channel, and domain, and tell the team where the corpus is failing them

Tooling & internal UIs

- Write the Python that holds all of this together - ingestion and processing scripts, validation checks, catalogue tooling, and reporting
- Build small internal web tools with modern coding assistants: annotation and QC interfaces, dataset explorers, audio A/B listening and diffing tools, and demo UIs for internal reviews and customer conversations
- Turn recurring manual work into tooling. If you have done it by hand three times, it should be a script or a screen by the fourth

Governance, security & customers

- Treat voice data as sensitive PII by default - voice recordings, transcripts, and speaker embeddings, the last of which are biometric data and carry the highest sensitivity
- Build de-identification and redaction into the pipeline: account and card numbers, Aadhaar and PAN, names, addresses, and health or financial detail out of transcripts before anyone works with them
- Maintain provenance and consent records well enough to survive an audit - what data came from where, under what licence or DPA, retained how long, and deleted on what trigger
- Apply data standards and recognised governance practice to how corpora are documented, accessed, retained, and disposed of - aligned with DPDP and, where relevant, RBI, IRDAI, and HIPAA expectations

Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.

📌 Data Engineer (Speech/Language Data) (Bengaluru)
🏢 Gnani Innovations
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: data engineer (speech/language data) (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: data engineer (speech/language data) (bengaluru) / bengaluru