08 Sep
|
Innodata
|
Sector 27
08 Sep
Innodata
Sector 27
ROLE OVERVIEW & OBJECTIVE
We are looking for exceptional Speech & Audio AI Evaluation Specialists with genuine in-house Global Capability Centre (GCC) / Captive international voice experience to evaluate state-of-the-art Speech-to-Speech (S2S), Text-to-Speech (TTS), and real-time conversational voice agents. Unlike traditional transcription or BPO roles, this specialist position demands a highly trained auditory ear to evaluate prosody, cadence, phonetic accuracy, emotional steering, and paralinguistic nuance across global English dialects. Candidates must possess C2 level near-native fluency to execute rigorous human-preference and synchronization benchmarks.
KEY RESPONSIBILITIES & CORE WORKFLOWS
•
S2S & TTS Naturalness Scoring:
Evaluate live Speech-to-Speech and neural Text-to-Speech outputs across intelligibility, rhythm, and conversational cadence. Paralinguistic & Emotion Steering Evaluation:
Benchmark how effectively models express nuanced vocal attributes including empathy, hesitation, tone inflection, irony, and situational urgency. Pairwise Audio Preference Ratings:
Conduct blinded, head-to-head auditory preference evaluations between candidate audio completions, providing detailed perceptual rationales. •
Multi-Modal data validations:
Verify multi-modal synchronization, lip-sync alignment, and acoustic scene consistency for video dubbing and avatar-driven speech models. Accent & Dialect Calibration:
Apply standardized rubric metrics uniformly across North American, British, Australian, and international English accents without regional bias. Defensible Auditory Documentation:
Document timestamped acoustic anomalies, unnatural vocal artifacts, metallic distortion, and hallucinated phonetic segments.
CANDIDATE PROFILE & QUALIFICATIONS
Mandatory Requirements Experience: 2+ years of dedicated international voice process experience exclusively within a Captive / In-House Global Capability Centre (GCC) serving US/global client bases. Language Mastery: C2 near-native English proficiency with absolute mastery over conversational nuances, colloquialisms, idioms, and tonal dynamics. Acoustic Acuity: Formally trained auditory ear for subtle vocal inflections, pitch modulation, cadence shifts, and articulation artifacts.
Preferred Qualifications
- Linguistics & Phonetics: Academic coursework or practical background in phonetics, phonology, auditory acoustics, or speech-language pathology.
- Speech QA Background: Prior qualified experience in voice quality analytics, acoustic data annotation, or speech synthesis evaluation.
- Acoustic Acuity: Formally trained auditory ear for subtle vocal inflections, pitch modulation, cadence shifts, and articulation artifacts.
- Multilingual Ability: Additional fluency in major European, Asian, or Latin American languages to support cross-lingual speech benchmarks.
📌 Analyst (Speech and Audio AI evaluation) (Sector 27)
🏢 Innodata
📍 Sector 27