22 Aug
|
Zilo AI
|
Hyderabad
Across every track, you will
- Drive AI research and innovation in Speech/Audio
AI
- Translate research questions into measurable hypotheses; design rigorous experiments and ablations, analyze model behavior and failure modes, reproduce and build on relevant papers
- Take research from prototype to production-grade, deployable systems
- Define appropriate objective and subjective evaluation metrics, and run benchmarking, error analysis, and statistical comparisons to know when a model is actually better – not just different
- Collaborate cross-functionally with engineering, product, and DSP/embedded teams
- Bring strong problem-solving ability – debug, iterate, and optimize models under real-world constraints
- Stay current with SOTA literature and bring current techniques into the team, publish or patent where appropriate
Track 1: Automatic Speech Recognition (ASR)
- Work on end-to-end ASR architectures (Transformer, Conformer, RNN-T) for streaming speech recognition
- Work with speech foundation models (Whisper-style architectures) and self-supervised models (HuBERT, wav2vec 2.0, WavLM) for representation learning and fine-tuning
- Own the end-to-end ASR pipeline – preprocessing, feature extraction, acoustic modeling, and decoding/language modeling
- Research streaming and causal architectures under strict latency constraints
- Improve robustness under noisy, multi-accent, and low-resource conditions – starting with Indian-English / regional Indian accents, with a roadmap to scale to international accents and languages
- Research accent-invariant / accent-normalized representations – removing accent information from learned embeddings to improve generalization across accents and geographies
- Study representation quality and robustness across speakers, accents, recording conditions, and domains
- Evaluated on: WER, CER, RTF, latency, memory
- Preferred background: prior work on CTC/RNN-T, self-supervised pretraining, or Whisper finetuning
Track 2: Text-to-Speech (TTS)
- Work on flow-matching / diffusion-based TTS systems (in the spirit of CosyVoice, ZipVoice, VITS, and similar SOTA architectures) for real-time, low-compute synthesis
- Experience with duration-controlled, time-aligned TTS using forced alignment (e.g., MFA) and Monotonic Alignment Search (MAS) for multi-speaker settings
- Research controllable accent adaptation using speaker, accent, linguistic, prosodic, and/or style representations – including decoder-side conditioning and representation-level approaches – starting with Indian-English and US-English accent pairs and scaling to international accents
- Research controllable duration, timing, rhythm, F0, and prosody modelling for natural accent rendering.
- Improve naturalness, prosody, and voice
-cloning quality
- Optimize inference for real-time, streaming synthesis on constrained hardware
- Evaluated on: MOS, speaker similarity, intelligibility, prosody, RTF
- Preferred background: prior work with neural vocoders, diffusion/flow-matching TTS, or zeroshot voice cloning
📌 Speech Language Specialist (Hyderabad)
🏢 Zilo AI
📍 Hyderabad