Data Engineer-Speech and language Data (Bengaluru)

Data Engineer-Speech and language Data (Bengaluru)

10 Sep
|
gnani.ai
|
Bengaluru

10 Sep

gnani.ai

Bengaluru

Data Engineer / Analyst — Speech & Language Data
www.gnani.ai
Location: Bengaluru Type: Full time Function: AI Team
About the role
Every speech model we ship is downstream of a decision someone made about data. Which hours got recorded. Which transcripts were trusted. Whether the Tamil test set accidentally shares speakers with the Tamil training set. Whether a number was written as २५ or 25. Those decisions set the ceiling on what our ASR and TTS models can ever achieve — and no amount of GPU time recovers a corpus that was built carelessly.
Training accurate ASR across 22+ Indian languages takes tens of thousands of hours of audio. Training TTS voices that sound human takes carefully recorded, tightly aligned, phonetically balanced speech. Right now that data lives across many places and many heads. We want one person to own it.
This role is the data owner for Gnani’s AI team. You source it, clean it, label it, version it, catalogue it, and defend its provenance — and you are the single point of contact when any researcher needs hours, a benchmark,



or an answer about where something came from. It is a hands-on engineering and analysis role with real judgment attached: you will listen to audio, read transcripts, and be the person who notices that something is wrong before a model trains on it.
Core mandate
Sourcing & acquisition
Find, evaluate, and bring in speech and text data across 22+ Indian languages — public, licensed, customer, and commissioned
Pipelines & data quality
Build the cleaning, segmentation, alignment, and validation pipelines that turn raw audio into trainable hours
Annotation & benchmarks
Own transcription and labelling quality, and maintain the frozen evaluation sets the whole team is judged on
Data ownership & governance
Be the single point of contact for data — catalogued, versioned, consented, and compliant by default

What you’ll drive
Sourcing & acquisition
Be permanently on the hunt for data — track new public Indic speech and text corpor

📌 Data Engineer-Speech and language Data (Bengaluru)
🏢 gnani.ai
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: data engineer-speech and language data (bengaluru) / bengaluru