Solid hands-on experience with CLIP and other contrastive embedding models, and Vision-Language Models (VLMs) for tasks such as captioning, visual QA, and grounding.
Practical experience designing Vision RAG systems, including multimodal embeddings, vector databases, and retrieval-augmented generation patterns.
Experience fine-tuning and adapting vision/VLM models using parameter-productive techniques (LoRA/QLoRA) and classical CV model training/transfer learning.
Working knowledge of hyperscaler AI/vision services across at least two of AWS, GCP, and Azure, and their trade-offs for vision workloads.
Familiarity with edge AI deployment hardware accelerators (NVIDIA Jetson, Coral/edge TPU), model compression, quantization, and runtime optimization.
Proficiency in Python and vision/ML frameworks (PyTorch, Hugging Face Transformers, OpenCV) and model format/runtime tooling (ONNX, TensorRT, OpenVINO).
Understanding of containerization and orchestration (Docker, Kubernetes) for scalable vision model serving.
Robust architectural and communication skills, with the ability to translate business needs into technical vision/GenAI solution designs.