05 Sep
|
Technozis
|
Mount Abu
05 Sep
Technozis
Mount Abu
Senior VLM Developer – Vision Language Models
Location: Bangalore, India
Experience: 7.5+ Years
Joining: Immediate – September 2026
We are looking for an experienced Senior VLM Developer with strong foundations in Machine Learning and Deep Learning to build, fine-tune, and evaluate Vision-Language Models (VLMs) for multimodal perception and understanding applications.
The role involves developing large-scale training pipelines, curating multimodal datasets, experimenting with model architectures, and optimizing models for downstream applications.
Key Responsibilities
- Design, implement, and fine-tune VLM architectures such as LLaVA, Qwen-VL, or similar models .
- Develop data preprocessing and training pipelines for large-scale multimodal datasets.
- Optimize model performance across visual-language benchmarks and internal use cases.
- Collaborate with research and synthetic data teams to integrate generated data into model training.
- Conduct ablation studies and experiments to evaluate model performance.
- Maintain documentation of experiments, model behavior, and evaluation results.
- Develop and optimize ML/DL solutions for complex business problems.
- Work independently on statistical, machine learning, and research-oriented projects.
- Collaborate with cross-functional teams to translate research and experimentation into practical AI solutions.
Required Skills
- 7.5+ years of experience in Machine Learning / Deep Learning / Data Science.
- Strong hands-on experience with Vision-Language Models (VLMs) and multimodal AI.
- Experience designing and fine-tuning VLM architectures such as LLaVA, Qwen-VL, or equivalent models .
- Strong hands-on expertise in PyTorch .
- Experience with PyTorch Lightning .
- Strong experience with the Hugging Face ecosystem.
- Strong understanding of Transformers and vision-language architectures .
- Experience with multimodal fusion techniques.
- Experience with distributed training frameworks .
- Hands-on experience with large-scale model training, fine-tuning, and evaluation.
- Experience with PEFT / parameter-efficient fine-tuning and/or reinforcement learning techniques.
- Robust experience in data preprocessing and multimodal dataset preparation .
- Experience working with synthetic data is preferred.
Mandatory Domain Experience
Candidates must have experience in at least one of the following domains:
CPG | Retail | Pharma
Candidates without relevant CPG, Retail, or Pharma domain experience should not be considered.
Education
B.Tech / M.Tech / Ph.D. in:
- Computer Science
- Artificial Intelligence
- Machine Learning
- Data Science
- Or a related technical discipline
Joining Requirement: Immediate joiners only.
Candidates must be able to join on or before 30 September 2026 .
Candidates with a joining date after 30 September 2026 will not be considered.
📌 Senior VLM Developer – Vision Language Models (Mount Abu)
🏢 Technozis
📍 Mount Abu