About Sarvam
- About the Role
You will work across the full lifecycle of vision-language model (VLM) development — data, training, evaluation, and production. The team's scope will evolve as the field does; we want researchers who are comfortable with that and can lead.
- What You'll Do
- Research vision-language architectures — encoders, fusion mechanisms, pretraining objectives, and scaling behaviour
- Design training methods (pretraining, SFT, RLHF, DPO) adapted for multilingual VLMs
- Investigate data strategies — what mixtures, quality signals, and synthetic data approaches actually move the needle
- Build evaluation frameworks and benchmarks, especially for Indic multimodal tasks
- Study model failure modes, robustness, and interpretability
- Work closely with engineers to ensure ideas are testable at scale — prototype fast, then validate properly
- Engage with the broader research community through open-source contributions and collaborations
- What We're Looking For
- Deep understanding of vision-language models — training dynamics, architecture tradeoffs,
and failure modes
- Track record of good research — through publications, technical reports, or impactful shipped work
- Rigorous experimental design — able to isolate variables and draw defensible conclusions
- Strong PyTorch skills — runs experiments end to end
- Intellectual range — willing to work across data, training, and evaluation problems
- Bonus Points
- PhD/Master's with relevant research experience in ML, Computer Vision, NLP, or related field
- Research papers published at A/A* venues
- Experience with multilingual or low-resource language modelling
- Familiarity with document understanding, OCR, or structured visual prediction
- Experience with large-scale data curation and its effect on model quality
- Why Sarvam?
Sarvam is a rapid-moving, high talent-density team building full-stack AI for India, working on problems that push the frontiers of AI with real population-scale impact.
- Work al
📌 Researcher, Vision (India)
🏢 SARVAM
📍 India