Senior AI Speech Engineer — Speech-to-Text (STT) & Text-to-Speech (TTS) Systems

Senior AI Speech Engineer — Speech-to-Text (STT) & Text-to-Speech (TTS) Systems

13 Sep
|
HYrEzy Tech Solutions
|
Gurugram

13 Sep

HYrEzy Tech Solutions

Gurugram

- Location: Hyderabad / NCR
- Employment Type: Full-Time
- Department: Core Speech AI &
- Audio Intelligence
- Experience Range: 5–9 Years

About The Company &

- Audio Engineering Culture

Voice is becoming the primary interaction layer for next-generation consumer and enterprise applications. As a high-growth AI technology enterprise, we build ultra-low latency, highly natural conversational voice AI agents, real-time transcription engines, and expressive voice synthesis systems across multiple Indian and global languages. Our audio AI engineering culture is built around advanced digital signal processing (DSP), neural vocoders, acoustic modeling, and real-time audio stream streaming. We train and deploy state-of-the-art speech models that operate under strict latency budgets for live voice interactions. If you are passionate about advanced audio machine learning and acoustic engineering, this is your arena.

Position Overview An expert, hands-on Senior AI Speech Engineer is needed to architect, train, and optimize our next-generation Speech-to-Text (Automatic Speech Recognition - ASR) and Text-to-Speech (Neural TTS &

- Voice Cloning) pipelines. In this role, you will own the end-to-end lifecycle of acoustic and language models designed to handle diverse accents, background noise, and real-time streaming audio.

You will collaborate closely with conversational AI product managers, backend streaming engineers, and linguistics specialists to build human-like voice interaction experiences.

Key Responsibilities &

- Technical Ownership1. Speech-to-Text (STT / ASR) Architecture

- Acoustic &
- Language Modeling: Train, fine-tune, and optimize end-to-end ASR architectures (Whisper, Conformer, CTC-RNN, Wav2Vec) for high-accuracy transcription in noisy environments and multi-lingual Indian contexts.
- Streaming &
- Real-Time Decoding:



Implement low-latency streaming recognition pipelines, custom vocabulary biasing, and real-time punctuation/inverse text normalization models.
- Acoustic Adaptation: Build speaker-adaptation and domain-specific fine-tuning workflows to improve word error rates (WER) for specialized enterprise jargon.
- Text-to-Speech (TTS) &
- Voice Synthesis
- Neural Voice Generation: Develop and optimize expressive TTS architectures (VITS, Tacotron 2, FastSpeech, XTTS) and neural vocoders (HiFi-GAN, WaveGlow) for hyper-realistic voice output.
- Voice Cloning &
- Prosody Control: Implement zero-shot voice cloning, emotion transfer, and fine-grained pitch/tempo control mechanisms.
- Streaming Synthesis: Build chunked, streaming text-to-speech pipelines that achieve sub-500ms time-to-first-audio (TTFA) metrics.
- Audio Processing &
- Edge Optimization
- Digital Signal Processing (DSP): Implement noise suppression, echo cancellation, acoustic beamforming, and voice activity detection (VAD) algorithms.
- Model Quantization &
- Deployment: Quantize and compile acoustic and speech models using ONNX Runtime, TensorRT, or CoreML for efficient cloud and edge execution.

Comprehensive Tech Stack &

- Technical RequirementsCore Technical Stack

- Languages: Advanced proficiency in Python and C++ (for high-performance audio processing wrappers).
- Deep Learning &
- Audio Frameworks: PyTorch, Kaldi, ESPnet, Hugging Face, Torchaudio, Librosa, Soundfile.
- Neural Vocoders &
- Architectures: HiFi-GAN,



WaveNet, Whisper, Conformer, FastSpeech.
- Serving &
- Streaming: NVIDIA Triton Inference Server, WebSockets, gRPC for low-latency audio streaming.
- Cloud &
- Infrastructure: AWS / GCP, Docker, Kubernetes, GPU optimization pipelines.

Experience &

- Educational Qualifications

- Experience: 5 to 9 years of professional software engineering experience, with at least 3+ years focused exclusively on speech technology, audio signal processing, ASR (STT), or TTS research and engineering.
- Education: Bachelor’s, Master’s, or Ph.D. degree in Computer Science, Electrical Engineering, Signal Processing, Acoustics, or a related discipline from a premier institution (IITs, IISc, or top global universities).

Competencies &
- Behavioral Traits

- Deep Domain Passion: Fascination with human phonetics, acoustic physics, psychoacoustics, and neural audio generation.
- Performance Obsession: Uncompromising focus on reducing audio latency, memory footprints, and word error rates in production environments.

Qualified Interview Process
- Initial Screening: Technical discussion covering signal processing fundamentals, ASR/TTS architectures, and past audio project scale.
- Audio System Design Round: Collaborative design exercise building a real-time conversational voice pipeline with sub-second latency.
- Coding &
- Deep Learning Assessment: Live coding session focusing on tensor manipulations for audio spectrograms or custom PyTorch training loops.
- Culture Fit &
- Leadership: Interview with engineering leadership focusing on innovation, execution velocity, and technical ownership.

Skills: speech-to-text (stt / asr) architecture,text-to-speech,models,signal processing,ttfa,coreml,tts,waveglow,onnx,speech,xtts,gan,dsp,tensorrt,processing,fastspeech

📌 Senior AI Speech Engineer — Speech-to-Text (STT) & Text-to-Speech (TTS) Systems
🏢 HYrEzy Tech Solutions
📍 Gurugram

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior ai speech engineer — speech-to-text (stt) & text-to-speech (tts) systems / gurugram

Subscribe to this job alert:

Get the latest job offers by email for: senior ai speech engineer — speech-to-text (stt) & text-to-speech (tts) systems / gurugram