Data Engineer-Speech and language Data (Bengaluru)

Data Engineer-Speech and language Data (Bengaluru)

09 Sep
|
gnani.ai
|
Bengaluru

09 Sep

gnani.ai

Bengaluru

Data Engineer / Analyst — Speech & Language Data www.gnani.ai

Location: Bengaluru Type: Full-time Function: AI Team

About the role

Every speech model we ship is downstream of a decision someone made about data. Which hours got recorded. Which transcripts were trusted. Whether the Tamil test set accidentally shares speakers with the Tamil training set. Whether a number was written as २५ or 25. Those decisions set the ceiling on what our ASR and TTS models can ever achieve — and no amount of GPU time recovers a corpus that was built carelessly.

Training accurate ASR across 22+ Indian languages takes tens of thousands of hours of audio. Training TTS voices that sound human takes carefully recorded, tightly aligned, phonetically balanced speech. Right now that data lives across many places and many heads. We want one person to own it.

This role is the data owner for Gnani’s AI team. You source it, clean it, label it, version it, catalogue it, and defend its provenance — and you are the single point of contact when any researcher needs hours, a benchmark, or an answer about where something came from. It is a hands-on engineering and analysis role with real judgment attached: you will listen to audio, read transcripts, and be the person who notices that something is wrong before a model trains on it.

Core mandate

Sourcing & acquisition

Find, evaluate, and bring in speech and text data across 22+ Indian languages — public, licensed, customer, and commissioned

Pipelines & data quality

Build the cleaning, segmentation, alignment, and validation pipelines that turn raw audio into trainable hours

Annotation & benchmarks

Own transcription and labelling quality, and maintain the frozen evaluation sets the whole team is judged on

Data ownership & governance

Be the single point of contact for data — catalogued, versioned, consented, and compliant by default

What you’ll drive

Sourcing & acquisition

- Be permanently on the hunt for data — track new public Indic speech and text corpora as they release, evaluate them for quality and licence terms, and bring the positive ones in fast
- Know the landscape cold: the AI4Bharat ecosystem (IndicVoices, Kathbath, Shrutilipi, GramVaani, IndicTTS), LIMMITS, SYSPIN, Rasa, Common Voice, FLEURS, and whatever lands next — and maintain an honest view of what each is actually good for
- Assess and onboard customer-contributed data — what consent covers it, what the DPA permits, what has to be de-identified before it can be used, and what cannot be used at all
- Spec and manage commissioned data collection and studio recording — script design for phonetic and dialect coverage, speaker recruitment criteria, session QC, and vendor throughput
- Keep a running gap analysis: hours per language, per dialect, per channel, per domain — so the team argues about priorities with numbers instead of impressions

Pipelines, cleaning & curation

- Build and own the pipelines that turn raw audio into trainable hours — format conversion and resampling, loudness normalisation, VAD-based segmentation, diarisation, and forced alignment
- Automate quality screening at scale: SNR estimation, clipping and DC offset, truncated utterances, silence-heavy segments, language-ID mismatches, and audio–transcript drift
- De-duplicate seriously — audio fingerprinting and near-duplicate text detection, so the same recording is not both trained on and evaluated on
- Own the text side, which for Indic data is most of the work:



orthographic convention, numeral form, transliteration and romanised code-switch, punctuation, casing, inverse text normalisation, and lexicon maintenance — inconsistent conventions show up as WER that has nothing to do with the model
- Run large-scale pseudo-labelling and weak supervision — transcribe unlabelled audio with existing models, filter on confidence and cross-model agreement, and route the uncertain cases to human review
- Version and catalogue everything: reproducible dataset snapshots, training manifests, dataset cards recording source, licence, consent basis, and known limitations

Annotation, listening & quality

- Own transcription and labelling quality end to end — write the annotation guidelines, train annotators against them, and revise them when reality disagrees
- Listen to audio and annotate yourself, as needed. Sampling real data is how you find the problems dashboards hide, and it is part of this job rather than beneath it
- Measure annotation quality properly — gold sets, inter-annotator agreement, WER between independent transcriptions — and manage vendors on quality and cost per hour, not volume alone
- Build the TTS-specific quality gates: alignment tightness, transcript fidelity, speaker and style consistency, prosody and emotion tagging, and codec round-trip checks
- Coordinate listening panels for subjective evaluation (MOS / CMOS / preference tests) — recruitment, screening, sample randomisation, and result analysis

Benchmarks & evaluation sets

- Own Gnani’s internal benchmark suite for ASR and TTS — curated, frozen, versioned, speaker-disjoint from training data, and representative across language, dialect, accent, channel, and domain
- Guarantee split hygiene. No test speaker in train, no leaked utterance, no silently mutated eval set — and be willing to block a result that violates it
- Track external and public benchmarks, reproduce them faithfully, and flag when a published comparison is not apples to apples
- Analyse results, not just produce them — slice error by language, speaker, channel, and domain, and tell the team where the corpus is failing them

Tooling & internal UIs

- Write the Python that holds all of this together — ingestion and processing scripts, validation checks, catalogue tooling, and reporting
- Build small internal web tools with modern coding assistants: annotation and QC interfaces, dataset explorers, audio A/B listening and diffing tools, and demo UIs for internal reviews and customer conversations
- Turn recurring manual work into tooling. If you have done it by hand three times, it should be a script or a screen by the fourth

Governance, security & customers

- Treat voice data as sensitive PII by default — voice recordings, transcripts, and speaker embeddings, the last of which are biometric data and carry the highest sensitivity
- Build de-identification and redaction into the pipeline: account and card numbers, Aadhaar and PAN, names, addresses, and health or financial detail out of transcripts before anyone works with them




- Maintain provenance and consent records well enough to survive an audit — what data came from where, under what licence or DPA, retained how long, and deleted on what trigger
- Apply data standards and recognised governance practice to how corpora are documented, accessed, retained, and disposed of — aligned with DPDP and, where relevant, RBI, IRDAI, and HIPAA expectations
- Work with the security team on access control, encryption at rest and in transit, least-privilege data rooms, and audit logging — and support customer security reviews and data-handling questions directly

Being the single point of contact

- Be the person the team comes to for data. Take requests from ASR, TTS, LLM, and product, understand what is actually needed, and deliver the right subset with the right metadata
- Communicate clearly and often across a team of researchers and engineers with different needs and different vocabularies, and keep everyone working from one shared picture of what data exists
- Say no when a request would compromise split hygiene, consent, or compliance — and explain why in terms the requester accepts

Who you are

EXPERIENCE WE’RE LOOKING FOR

- 3–6 years in data engineering, data analysis, or ML data work — with real ownership of a dataset that other people depended on
- Strong, practical Python: pandas or polars, scripting, automation, and comfort working with files and pipelines at scale rather than only in notebooks
- Genuine ownership instinct. This role has no one above it to catch a data problem — you are that person, and you should want to be
- Excellent written and verbal communication, and the patience to work across several stakeholders with competing data needs
- Careful, detail-obsessed temperament — the kind of person who spots that two files have subtly different transcript conventions before anyone trains on them

WHAT MAKES A STANDOUT CANDIDATE

- Hands-on experience with audio or speech data — segmentation, alignment, transcription workflows, or annotation operations
- Fluency in one or more Indian languages beyond English, and an ear for dialect and code-switching
- Experience running annotation vendors or an in-house labelling team, including quality measurement and cost management
- Comfort building small web UIs and demos — Streamlit, Gradio, or a lightweight FastAPI plus React app, coding assistants very much welcome
- Exposure to data governance and privacy in a regulated setting — DPDP, consent management, PII redaction, or supporting customer security reviews
- Comfortable operating with high ownership in a fast-moving, post–Series B environment

TECHNICAL FLUENCY — A MUST-HAVE

- Python & data: pandas / polars, PyArrow and Parquet, JSONL manifests, SQL, and clean reusable scripting
- Audio tooling: ffmpeg, sox, librosa or torchaudio — resampling, segmentation, loudness, and basic signal sanity checks
- Pipelines & storage: object storage, sharded formats for training throughput, workflow orchestration (Airflow, Prefect, or Dagster), Git, Docker
- Dataset discipline: versioning and snapshots, train / dev / test split design, speaker-disjoint splits, deduplication, dataset documentation
- Analysis & reporting: WER and CER computation and error slicing, coverage and gap reporting, agreement metrics, and dashboards people actually read
- Nice to have: NVIDIA NeMo manifest conventions, forced alignment tools, Spark or Ray for large jobs, and basic familiarity with how ASR and TTS models consume data

📌 Data Engineer-Speech and language Data (Bengaluru)
🏢 gnani.ai
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: data engineer-speech and language data (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: data engineer-speech and language data (bengaluru) / bengaluru