Senior VLM Developer – Vision Language Models Location: Bangalore, India Experience: 7.5 Years Joining: Immediate – September 2026 We are looking for an experienced Senior VLM Developer with strong foundations in Machine Learning and Deep Learning to build, fine-tune, and evaluate Vision-Language Models (VLMs) for multimodal perception and understanding applications. The role involves developing large-scale training pipelines, curating multimodal datasets, experimenting with model architectures, and optimizing models for downstream applications. Key Responsibilities Design, implement, and fine-tune VLM architectures such as LLaVA, Qwen-VL, or similar models . Develop data preprocessing and training pipelines for large-scale multimodal datasets. Optimize model performance across visual-language benchmarks and internal use cases. Collaborate with research and synthetic data teams to integrate generated data into model training. Conduct ablation studies and experiments to evaluate model performance. Maintain documentation of experiments, model behavior, and evaluation results. Develop and optimize ML/DL solutions for complex business problems. Work independently on statistical, machine learning, and research-oriented projects. Collaborate with cross-functional teams to translate research and experimentation into practical AI solutions. Required Skills 7.5 years of experience in Machine Learning / Deep Learning / Data Science.
Strong hands-on experience with Vision-Language Models (VLMs) and multimodal AI. Experience designing and fine-tuning VLM architectures such as LLaVA, Qwen-VL, or equivalent models . Strong hands-on expertise in PyTorch . Experience with PyTorch Lightning . Strong experience with the Hugging Face ecosystem. Solid understanding of Transformers and vision-language architectures . Experience with multimodal fusion techniques. Experience with distributed training frameworks . Hands-on experience with large-scale model training, fine-tuning, and evaluation. Experience with PEFT / parameter-efficient fine-tuning and/or reinforcement learning techniques. Strong experience in data preprocessing and multimodal dataset preparation . Experience working with synthetic data is preferred. Mandatory Domain Experience Candidates must have experience in at least one of the following domains: CPG | Retail | Pharma Candidates without relevant CPG, Retail, or Pharma domain experience should not be considered. Education B.Tech / M.Tech / Ph.D. in: Computer Science Artificial Intelligence Machine Learning Data Science Or a related technical discipline Joining Requirement: Immediate joiners only. Candidates must be able to join on or before 30 September 2026 . Candidates with a joining date after 30 September 2026 will not be considered.
📌 Senior VLM Developer – Vision Language Models (India)
🏢 Technozis
📍 India
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.