08 Sep
|
Technozis
|
Bengaluru
08 Sep
Technozis
Bengaluru
Senior VLM Developer – Vision Language Models
nLocation: Bangalore, India
nExperience: 7.5+ Years
nJoining: Immediate – September 2026
n
nWe are looking for an experienced Senior VLM Developer with strong foundations in Machine Learning and Deep Learning to build, fine-tune, and evaluate Vision-Language Models (VLMs) for multimodal perception and understanding applications.
n
nThe role involves developing large-scale training pipelines, curating multimodal datasets, experimenting with model architectures, and optimizing models for downstream applications.
n
nKey Responsibilities
n
n
- Design, implement, and fine-tune VLM architectures such as LLaVA, Qwen-VL, or similar models.
n
- Develop data preprocessing and training pipelines for large-scale multimodal datasets.
n
- Optimize model performance across visual-language benchmarks and internal use cases.
n
- Collaborate with research and synthetic data teams to integrate generated data into model training.
n
- Conduct ablation studies and experiments to evaluate model performance.
n
- Maintain documentation of experiments, model behavior, and evaluation results.
n
- Develop and optimize ML/DL solutions for complex business problems.
n
- Work independently on statistical, machine learning, and research-oriented projects.
n
- Collaborate with cross-functional teams to translate research and experimentation into practical AI solutions.
n
n
nRequired Skills
n
n
- 7.5+ years of experience in Machine Learning / Deep Learning / Data Science.
n
- Strong hands-on experience with Vision-Language Models (VLMs) and multimodal AI.
n
- Experience designing and fine-tuning VLM architectures such as LLaVA, Qwen-VL, or equivalent models.
n
- Strong hands-on expertise in PyTorch.
n
- Experience with PyTorch Lightning.
n
- Strong experience with the Hugging Face ecosystem.
n
- Robust understanding of Transformers and vision-language architectures.
n
- Experience with multimodal fusion techniques.
n
- Experience with distributed training frameworks.
n
- Hands-on experience with large-scale model training, fine-tuning, and evaluation.
n
- Experience with PEFT / parameter-efficient fine-tuning and/or reinforcement learning techniques.
n
- Strong experience in data preprocessing and multimodal dataset preparation.
n
- Experience working with synthetic data is preferred.
n
n
nMandatory Domain Experience
nCandidates must have experience in at least one of the following domains:
nCPG | Retail | Pharma
nCandidates without relevant CPG, Retail, or Pharma domain experience should not be considered.
n
nEducation
nB.Tech / M.Tech / Ph.D. in:
n
n
- Computer Science
n
- Artificial Intelligence
n
- Machine Learning
n
- Data Science
n
- Or a related technical discipline
n
n
nJoining Requirement: Immediate joiners only.
nCandidates must be able to join on or before 30 September 2026.
nCandidates with a joining date after 30 September 2026 will not be considered.
📌 Senior VLM Developer - Vision Language Models (Bengaluru)
🏢 Technozis
📍 Bengaluru