Senior AI Engineer – Vision Language Models (VLM) | Experience: 4+ Years Relevant Experience: 2+ Years in Multimodal AI / VLM | Contract | Work Mode: Remote (India)

Senior AI Engineer – Vision Language Models (VLM) | Experience: 4+ Years Relevant Experience: 2+ Years in Multimodal AI / VLM | Contract | Work Mode: Remote (India)

14 Aug
|
Unicorn workforce
|
India

14 Aug

Unicorn workforce

India

Position: Senior AI Engineer – Vision Language Models (VLM)

Experience: 4+ Years

Relevant Experience: 2+ Years in Multimodal AI / VLM

Employment Type: Contractual

Work Mode: Remote

Notice Period: Immediate Joiners / Short Notice Preferred

About the Role

We are looking for a Senior AI Engineer – Vision Language Models (VLM) with strong hands-on experience in Multimodal AI, Computer Vision, Generative AI, and Video Analytics .

The selected candidate will design, develop, optimize, and deploy AI solutions focused on video understanding, scene analysis, activity recognition, and real-time incident detection .

The role requires experience taking AI solutions from research and prototyping through production deployment , with strong expertise in VLMs, deep learning, Python, PyTorch, and GPU-based inference.

Key Responsibilities

VLM & Multimodal AI

- Design and develop Vision Language Model (VLM) solutions for video understanding and incident detection.
- Build multimodal AI systems combining visual, textual, and contextual information.
- Work with modern multimodal foundation models such as GPT-4o Vision, Gemini, Qwen-VL, LLaVA, InternVL, Florence , or equivalent.
- Develop prompting, reasoning, and evaluation strategies for complex visual scenarios.
- Improve model performance through experimentation, prompt engineering, feedback loops, and iterative optimization.

Video Understanding & Incident Detection
- Develop AI pipelines for real-time or near-real-time video analysis .
- Build solutions for detecting incidents such as:
- Falls
- Physical restraints
- Altercations
- Unsafe or unusual activities
- Other safety-related incidents
- Work on scene understanding,



activity recognition, action recognition, temporal reasoning, and video analytics .
- Process and analyze large-scale image and video datasets.

Model Evaluation & Optimization
- Benchmark VLMs against traditional Computer Vision and Action Recognition models .
- Define appropriate evaluation methodologies and performance metrics.
- Measure and optimize accuracy, precision, recall, F1 score, latency, and inference performance .
- Conduct experiments to identify the most effective models and approaches for specific use cases.

Production Deployment
- Deploy and optimize AI models across cloud and/or edge GPU infrastructure .
- Build productive inference pipelines for real-time video processing.
- Optimize models for GPU utilization, latency, scalability, and resource efficiency .
- Containerize AI applications using Docker.
- Apply MLOps best practices for model deployment and lifecycle management.
- Troubleshoot AI models and production inference pipelines.

Collaboration
- Collaborate with Data Scientists, ML Engineers, Software Engineers, and Product teams .
- Translate business requirements into scalable AI solutions.
- Contribute to technical architecture, design, testing, documentation, and optimization.
- Take solutions from POC/research stage to production .





Mandatory Technical Skills

AI / ML

- Strong proficiency in Python .
- Hands-on experience with PyTorch .
- Strong understanding of:
- Computer Vision
- Deep Learning
- Generative AI
- Multimodal AI
- Vision Language Models

Vision Language Models Hands-on experience with one or more of:
- GPT-4o Vision
- Google Gemini
- Qwen-VL
- LLaVA
- InternVL
- Florence
- Equivalent VLM / Multimodal Foundation Models

Computer Vision & Video Analytics
- Video analytics
- Scene understanding
- Activity recognition
- Action recognition
- Image/video processing
- Temporal reasoning
- Large-scale image and video datasets

Tools & Platforms
- Hugging Face
- OpenCV
- Docker
- REST APIs
- Azure AI/ML
- GPU-based model deployment
- MLOps

GPU & Model Deployment
- GPU optimization and efficient model inference.
- Experience deploying AI/ML models on cloud or edge GPU infrastructure .
- Understanding of inference latency, scalability, resource utilization, and model optimization.

Qualifications
- Bachelor's or Master's degree in Computer Science, Artificial Intelligence, Machine Learning, Data Science, Computer Engineering , or a related field.
- 4+ years of professional AI/ML, Computer Vision, Applied AI, or related experience.
- 2+ years of hands-on Multimodal AI / VLM experience.
- Proven experience taking AI solutions from research/prototyping to production .
- Strong understanding of AI evaluation metrics including Accuracy, Precision, Recall, F1 Score, and Latency .
- Experience working with large-scale image/video datasets.
- Strong analytical and problem-solving skills.

📌 Senior AI Engineer – Vision Language Models (VLM) | Experience: 4+ Years Relevant Experience: 2+ Years in Multimodal AI / VLM | Contract | Work Mode: Remote (India)
🏢 Unicorn workforce
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior ai engineer – vision language models (vlm) | experience: 4+ years relevant experience: 2+ years in multimodal ai / vlm | contract | work mode: remote (india) / india

Subscribe to this job alert:

Get the latest job offers by email for: senior ai engineer – vision language models (vlm) | experience: 4+ years relevant experience: 2+ years in multimodal ai / vlm | contract | work mode: remote (india) / india