15 Sep
|
Skyleaf Consultants LLP
|
Delhi
15 Sep
Skyleaf Consultants LLP
Delhi
About the Role
We are looking for a talented Multimodal AI Engineer to design, develop, and deploy AI systems that understand and reason across multiple modalities, including text, images, audio, video, and structured data .
You will work on the development of multimodal AI applications using Large Language Models (LLMs), Vision-Language Models (VLMs), speech models, embeddings, retrieval systems, and agentic architectures . The ideal candidate combines robust machine-learning fundamentals with hands-on experience building production-grade AI systems.
Key Responsibilities
- Design and develop multimodal AI solutions combining text, vision, audio, and video.
- Build applications using LLMs, VLMs, multimodal foundation models, embeddings, and generative AI models .
- Develop pipelines for processing and understanding images, documents, audio, video, and other unstructured data.
- Fine-tune, evaluate, and optimize foundation models for specific business use cases.
- Implement RAG, multimodal RAG, semantic search, vector databases, and knowledge-grounded generation .
- Develop AI agents and workflows capable of reasoning over multiple modalities.
- Work with APIs and open-source models such as GPT, Gemini, Claude, Llama, Qwen, Mistral, or similar models .
- Build model evaluation frameworks covering accuracy, hallucination, robustness, latency, and cost.
- Optimize inference performance and model serving for production workloads.
- Collaborate with data scientists, ML engineers, software engineers, and product teams to take AI solutions from prototype to production.
- Design scalable data and model pipelines for training, fine-tuning, evaluation, and inference.
- Stay current with advances in multimodal foundation models, computer vision, speech AI, generative AI, and agentic systems.
Required Skills & Experience
- 3+ years of experience in Machine Learning / Deep Learning / AI engineering .
- Strong programming skills in Python and experience with production software development.
- Strong understanding of transformers, attention mechanisms, embeddings, and deep-learning architectures .
- Hands-on experience with one or more of:
- Large Language Models (LLMs)
- Vision-Language Models (VLMs)
- Computer Vision
- Speech / Audio AI
- Video understanding
- Multimodal generative AI
- Experience with PyTorch or TensorFlow.
- Experience working with model APIs and/or open-source foundation models.
- Experience with RAG, vector databases, embeddings, and prompt engineering .
- Familiarity with model fine-tuning techniques such as LoRA, QLoRA, PEFT, or supervised fine-tuning .
- Understanding of ML evaluation, experimentation, and performance optimization.
- Experience deploying AI/ML systems using cloud platforms and containerized environments.
Good to Have
- Experience with multimodal models such as GPT-4o/5-class models, Gemini, Claude, Llama, Qwen-VL, or similar models .
- Experience with OCR, document intelligence, image understanding, video analytics, or speech recognition .
- Experience building multimodal RAG systems.
- Experience with agent frameworks and tool-calling architectures.
- Familiarity with LangChain, LlamaIndex, Hugging Face Transformers, vLLM, or similar frameworks .
- Experience with vector databases such as Pinecone, Milvus, Weaviate, Qdrant, or FAISS .
- Experience with AWS, Azure, or GCP.
- Knowledge of GPU optimization, quantization, distributed inference, or model serving.
- Experience with MLOps, CI/CD, Docker, Kubernetes, and model monitoring.
- Contributions to open-source projects, research publications, patents, or strong AI/ML side projects.
Education
- Bachelor's or Master's degree in Computer Science, Artificial Intelligence, Machine Learning, Electrical Engineering, Mathematics , or a related field.
- Equivalent practical experience and a strong portfolio of AI projects will also be considered.
📌 Multimodal AI Engineer (Delhi)
🏢 Skyleaf Consultants LLP
📍 Delhi