14 Aug
|
Unicorn workforce
|
India
14 Aug
Unicorn workforce
India
Position: Senior AI Engineer – Vision Language Models (VLM)
Experience: 4+ Years
Relevant Experience: 2+ Years in Multimodal AI / VLM
Employment Type: Contractual
Work Mode: Remote
Notice Period: Immediate Joiners / Short Notice Preferred
About the Role
We are looking for a Senior AI Engineer – Vision Language Models (VLM) with strong hands-on experience in Multimodal AI, Computer Vision, Generative AI, and Video Analytics .
The selected candidate will design, develop, optimize, and deploy AI solutions focused on video understanding, scene analysis, activity recognition, and real-time incident detection .
The role requires experience taking AI solutions from research and prototyping through production deployment , with strong expertise in VLMs, deep learning, Python, PyTorch, and GPU-based inference.
Key Responsibilities
VLM & Multimodal AI
- Design and develop Vision Language Model (VLM) solutions for video understanding and incident detection.
- Build multimodal AI systems combining visual, textual, and contextual information.
- Work with modern multimodal foundation models such as GPT-4o Vision, Gemini, Qwen-VL, LLaVA, InternVL, Florence , or equivalent.
- Develop prompting, reasoning, and evaluation strategies for complex visual scenarios.
- Improve model performance through experimentation, prompt engineering, feedback loops, and iterative optimization.
Video Understanding & Incident Detection
- Develop AI pipelines for real-time or near-real-time video analysis .
- Build solutions for detecting incidents such as:
- Falls
- Physical restraints
- Altercations
- Unsafe or unusual activities
- Other safety-related incidents
- Work on scene understanding,
activity recognition, action recognition, temporal reasoning, and video analytics .
- Process and analyze large-scale image and video datasets.
Model Evaluation & Optimization
- Benchmark VLMs against traditional Computer Vision and Action Recognition models .
- Define appropriate evaluation methodologies and performance metrics.
- Measure and optimize accuracy, precision, recall, F1 score, latency, and inference performance .
- Conduct experiments to identify the most effective models and approaches for specific use cases.
Production Deployment
- Deploy and optimize AI models across cloud and/or edge GPU infrastructure .
- Build productive inference pipelines for real-time video processing.
- Optimize models for GPU utilization, latency, scalability, and resource efficiency .
- Containerize AI applications using Docker.
- Apply MLOps best practices for model deployment and lifecycle management.
- Troubleshoot AI models and production inference pipelines.
Collaboration
- Collaborate with Data Scientists, ML Engineers, Software Engineers, and Product teams .
- Translate business requirements into scalable AI solutions.
- Contribute to technical architecture, design, testing, documentation, and optimization.
- Take solutions from POC/research stage to production .
Mandatory Technical Skills
AI / ML
- Strong proficiency in Python .
- Hands-on experience with PyTorch .
- Strong understanding of:
- Computer Vision
- Deep Learning
- Generative AI
- Multimodal AI
- Vision Language Models
Vision Language Models Hands-on experience with one or more of:
- GPT-4o Vision
- Google Gemini
- Qwen-VL
- LLaVA
- InternVL
- Florence
- Equivalent VLM / Multimodal Foundation Models
Computer Vision & Video Analytics
- Video analytics
- Scene understanding
- Activity recognition
- Action recognition
- Image/video processing
- Temporal reasoning
- Large-scale image and video datasets
Tools & Platforms
- Hugging Face
- OpenCV
- Docker
- REST APIs
- Azure AI/ML
- GPU-based model deployment
- MLOps
GPU & Model Deployment
- GPU optimization and efficient model inference.
- Experience deploying AI/ML models on cloud or edge GPU infrastructure .
- Understanding of inference latency, scalability, resource utilization, and model optimization.
Qualifications
- Bachelor's or Master's degree in Computer Science, Artificial Intelligence, Machine Learning, Data Science, Computer Engineering , or a related field.
- 4+ years of professional AI/ML, Computer Vision, Applied AI, or related experience.
- 2+ years of hands-on Multimodal AI / VLM experience.
- Proven experience taking AI solutions from research/prototyping to production .
- Strong understanding of AI evaluation metrics including Accuracy, Precision, Recall, F1 Score, and Latency .
- Experience working with large-scale image/video datasets.
- Strong analytical and problem-solving skills.
📌 Senior AI Engineer – Vision Language Models (VLM) | Experience: 4+ Years Relevant Experience: 2+ Years in Multimodal AI / VLM | Contract | Work Mode: Remote (India)
🏢 Unicorn workforce
📍 India