Senior VLM Developer – Vision Language Models (Bengaluru)

Senior VLM Developer – Vision Language Models (Bengaluru)

05 Sep
|
Technozis
|
Bengaluru

05 Sep

Technozis

Bengaluru

Senior VLM Developer – Vision Language Models

Location: Bangalore, India

Experience: 7.5+ Years

Joining: Immediate – September 2026

We are looking for an experienced Senior VLM Developer with strong foundations in Machine Learning and Deep Learning to build, fine-tune, and evaluate Vision-Language Models (VLMs) for multimodal perception and understanding applications.

The role involves developing large-scale training pipelines, curating multimodal datasets, experimenting with model architectures, and optimizing models for downstream applications.

Key Responsibilities

- Design, implement, and fine-tune VLM architectures such as LLaVA, Qwen-VL, or similar models.
- Develop data preprocessing and training pipelines for large-scale multimodal datasets.
- Optimize model performance across visual-language benchmarks and internal use cases.
- Collaborate with research and synthetic data teams to integrate generated data into model training.
- Conduct ablation studies and experiments to evaluate model performance.
- Maintain documentation of experiments, model behavior, and evaluation results.
- Develop and optimize ML/DL solutions for complex business problems.
- Work independently on statistical, machine learning, and research-oriented projects.
- Collaborate with cross-functional teams to translate research and experimentation into practical AI solutions.

Required Skills

- 7.5+ years of experience in Machine Learning / Deep Learning / Data Science.




- Strong hands-on experience with Vision-Language Models (VLMs) and multimodal AI.
- Experience designing and fine-tuning VLM architectures such as LLaVA, Qwen-VL, or equivalent models.
- Strong hands-on expertise in PyTorch.
- Experience with PyTorch Lightning.
- Strong experience with the Hugging Face ecosystem.
- Strong understanding of Transformers and vision-language architectures.
- Experience with multimodal fusion techniques.
- Experience with distributed training frameworks.
- Hands-on experience with large-scale model training, fine-tuning, and evaluation.
- Experience with PEFT / parameter-effective fine-tuning and/or reinforcement learning techniques.
- Strong experience in data preprocessing and multimodal dataset preparation.
- Experience working with synthetic data is preferred.

Mandatory Domain Experience

Candidates must have experience in at least one of the following domains:

CPG | Retail | Pharma

Candidates without relevant CPG, Retail, or Pharma domain experience should not be considered.

Education

B.Tech / M.Tech / Ph.D. in:

- Computer Science
- Artificial Intelligence
- Machine Learning
- Data Science
- Or a related technical discipline

Joining Requirement: Immediate joiners only.

Candidates must be able to join on or before 30 September 2026.

Candidates with a joining date after 30 September 2026 will not be considered.

📌 Senior VLM Developer – Vision Language Models (Bengaluru)
🏢 Technozis
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior vlm developer – vision language models (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: senior vlm developer – vision language models (bengaluru) / bengaluru