05 Sep
|
Technozis
|
Bengaluru
05 Sep
Technozis
Bengaluru
Senior VLM Developer – Vision Language Models
Location: Bangalore, India
Experience: 7.5+ Years
Joining: Immediate – September 2026
We are looking for an experienced Senior VLM Developer with strong foundations in Machine Learning and Deep Learning to build, fine-tune, and evaluate Vision-Language Models (VLMs) for multimodal perception and understanding applications.
The role involves developing large-scale training pipelines, curating multimodal datasets, experimenting with model architectures, and optimizing models for downstream applications.
Key Responsibilities
Design, implement, and fine-tune VLM architectures such as LLaVA, Qwen-VL, or similar models .
Develop data preprocessing and training pipelines for large-scale multimodal datasets.
Optimize model performance across visual-language benchmarks and internal use cases.
Collaborate with research and synthetic data teams to integrate generated data into model training.
Conduct ablation studies and experiments to evaluate model performance.
Maintain documentation of experiments, model behavior, and evaluation results.
Develop and optimize ML/DL solutions for complex business problems.
Work independently on statistical, machine learning, and research-oriented projects.
Collaborate with cross-functional teams to translate research and experimentation into practical AI solutions.
Required Skills
7.5+ years of experience in Machine Learning / Deep Learning / Data Science.
Strong hands-on experience with Vision-Language Models (VLMs) and multimodal AI.
Experience designing and fine-tuning VLM architectures such as LLaVA, Qwen-VL, or equivalent models .
Strong hands-on expertise in PyTorch .
Experience with PyTorch Lightning .
Robust experience with the Hugging Face ecosystem.
Strong understanding of Transformers and vision-language architectures .
Experience with multimodal fusion techniques.
Experience with distributed training frameworks .
Hands-on experience with large-scale model training, fine-tuning, and evaluation.
Experience with PEFT / parameter-efficient fine-tuning and/or reinforcement learning techniques.
Strong experience in data preprocessing and multimodal dataset preparation .
Experience working with synthetic data is preferred.
Mandatory Domain Experience
Candidates must have experience in at least one of the following domains:
CPG | Retail | Pharma
Candidates without relevant CPG, Retail, or Pharma domain experience should not be considered.
Education
B.Tech / M.Tech / Ph.D. in:
Computer Science
Artificial Intelligence
Machine Learning
Data Science
Or a related technical discipline
Joining Requirement: Immediate joiners only.
Candidates must be able to join on or before 30 September 2026 .
Candidates with a joining date after 30 September 2026 will not be considered.
📌 Senior VLM Developer – Vision Language Models (Bengaluru)
🏢 Technozis
📍 Bengaluru