06 Aug
|
Josh Talks
|
Gurugram
06 Aug
Josh Talks
Gurugram
Backend Engineer Intern
Location: Gurgaon, India (On-site Role)
Type: Full-time Internship (6 - 12 Months) | Paid
Eligibility: Students graduating in 2027 (B.Tech / B.E. — CSE, EE, AI/ML, or related) About JoshTalks AI
At JoshTalks AI, we believe voice will become the primary interface between humans and machines. We build benchmarks, datasets, and AI systems that power some of the world's most widely used speech models.
Our mission
- Enable machines to communicate as naturally as humans
- Build benchmarks and datasets that form the backbone of global speech AI progress
- Drive improvements through high-quality, diverse, real-world data — not just compute What You'll Work On
This is not a conventional internship. You will directly contribute to work that influences the global speech AI ecosystem.
Data Delivery (Cloud & Containers)
- Deliver processed speech and audio datasets via cloud-based container pipelines (Docker, Kubernetes)
- Design and manage end-to-end data delivery workflows from raw ingestion to structured output
- Build and maintain ETL pipelines to prepare data for model training and evaluation at scale
- Ensure reliable, versioned, and reproducible data handoffs to ML training infrastructure
- Monitor pipeline health, resolve bottlenecks, and optimize throughput for large-scale data movement Benchmarking & Evaluation
- Design and run evaluations for ASR and speech-to-speech systems
- Benchmark leading speech models to identify real-world strengths, weaknesses, and failure modes
- Build evaluation frameworks that guide top AI labs on model performance and improvement areas Modeling & Fine-Tuning
- Fine-tune speech recognition models (e.g., Whisper, wav2vec2) targeting Word Error Rates of ~5%
- Experiment with multilingual, code-switched, accented,
and noisy speech data
- Work with large-scale, production-grade speech datasets across 20+ Indian languages Who We're Looking For
Required Qualifications
- Students in t B.Tech/B.E. graduating in 2027 (CSE, EE, AI/ML, or related fields)
- Strong interest in speech, audio, NLP, or multimodal AI
- Hands-on experience in one or more of the following:
◦ Fine-tuning speech or language models (Whisper, wav2vec2, HuBERT, etc.)
◦ Building speech-driven pipelines, classifiers, or assistants
◦ Working with PyTorch, TensorFlow, or Hugging Face Transformers
◦ Cloud platforms (AWS, GCP, Azure) and container tools (Docker, Kubernetes) Preferred Qualifications
- Open-source contributions, GitHub projects, Kaggle experience,
- Experience with multilingual or low-resource speech data
- Familiarity with ETL workflows, data versioning (DVC, MLflow), or orchestration (Airflow, Prefect) Why Join Us
Meaningful Ownership
Own problems of global relevance from reducing ASR error rates to delivering data that trains the next generation of speech models.
Front-Row Seat to Speech AI
Your pipelines and datasets will directly influence benchmarks used by the world's leading AI research labs.
Deep Technical Learning
Work alongside experts building systems across 20+ Indian languages and real-world audio conditions.
Startup Environment, Global Reach A small, focused team working on problems that impact billions of users worldwide. How to Apply
If you are passionate about making speech AI as natural as human conversation and want to work at the true frontier of the field, we would love to hear from you.
Send your profile to: Note: This is a paid, full time, in-office internship based in Gurgaon. Remote work is not available for this role.
📌 Back End Developer Intern (Gurugram)
🏢 Josh Talks
📍 Gurugram