Push the frontier of multimodal video understanding, vision, audio, speech, and text fused into a single model.
What you'll do
Design and run experiments on large-scale multimodal models.
Publish or open-source impactful results where appropriate.
Collaborate with engineering to ship research into production.
What we're looking for
PhD (or equivalent experience) in ML, CV, NLP, or speech.
Track record of robust publications or production-shipped models.
Deep experience with large-scale model training.
Nice to have
Experience with video foundation models or long-context architectures.
Apply for this role
Tell us about yourself, it takes about 2 minutes.
First name
Last name
Email
Phone number
LinkedIn profile (optional)
Resume link (Google Drive, Dropbox, or personal site)
Portfolio / GitHub (optional)
Why this role?
By clicking submit, you agree to Zyris's Privacy Policy.
📌 Research Scientist, Multimodal (Bengaluru)
🏢 Zyris AI
📍 Bengaluru
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.