04 Aug
|
iMerit Technology
|
Bengaluru
04 Aug
iMerit Technology
Bengaluru
Machine Learning Engineer – Multimodal AI & Agentic systems
Location: Bengaluru, India
Experience: 2–4 Years (Relevant Industry Experience)
About iMerit iMerit is a global leader in AI data solutions, enabling organizations to build reliable, production-grade Artificial Intelligence systems through high-quality data, advanced tooling, and human expertise. We partner with leading enterprises in autonomous mobility, healthcare, geospatial intelligence, robotics, and generative AI to solve complex machine learning challenges.
With over 5,500 professionals worldwide, iMerit combines cutting-edge technology with human intelligence to accelerate the development and deployment of AI systems while maintaining the highest standards of quality, security, and social impact.
About the Role
We are looking for a passionate Machine Learning Engineer – Multimodal AI & Agentic systems to join our AI Engineering team. The role combines applied research, machine learning engineering, and platform development to build production-grade systems that understand, process, and reason across multiple data modalities, including images, video, text, documents, point clouds, and sensor data.
You will develop multimodal machine learning solutions and agentic systems for data curation, intelligent annotation, model evaluation, quality validation, synthetic data generation, and automated AI workflows. The role involves applying computer vision models, Large Language Models, Vision Language Models, multimodal foundation models, and tool-using agents to solve complex real-world problems.
The ideal candidate enjoys conducting structured experiments, translating research into reliable engineering systems, and building AI workflows with appropriate evaluation, observability, safeguards, and human oversight
Key Responsibilities
Computer Vision & Machine Learning
- Design, develop, train, and optimize state-of-the-art computer vision models for tasks including object detection,
semantic and instance segmentation, object tracking, pose estimation, OCR, classification, and 3D perception.
- Develop perception algorithms utilizing camera, LiDAR, radar, and multimodal sensor data.
- Build scalable training and inference pipelines for production deployment.
- Improve model accuracy through data-centric AI techniques including active learning, hard example mining, edge-case discovery, and human-in-the-loop learning.
- Evaluate and benchmark models using industry-standard metrics and datasets.
Agentic AI & Generative AI
- Design and develop Agentic AI systems capable of autonomous planning, reasoning, tool orchestration, and workflow execution.
- Build intelligent AI agents using modern LLM frameworks and orchestration platforms.
- Develop solutions for LLM evaluation, prompt engineering, synthetic data generation, knowledge retrieval, and autonomous data processing.
- Fine-tune, evaluate, and optimize open-source and proprietary Large Language Models and Small Language Models (SLMs).
- Integrate Vision Language Models (VLMs) into computer vision workflows for multimodal reasoning and data quality improvement.
AI Platform & Product Engineering
- Design modular, reusable AI services that can be integrated into enterprise products.
- Build APIs and microservices for AI inference and model serving.
- Develop automated validation, testing, benchmarking, and monitoring pipelines for machine learning systems.
- Collaborate with MLOps teams to deploy scalable AI services on cloud and edge infrastructure.
- Contribute to system architecture, technical design documents,
and engineering best practices.
Collaboration
- Work closely with Product Managers, Data Scientists, Annotation Experts, and Software Engineers to translate business requirements into AI solutions.
- Participate in technical design discussions, code reviews, and architecture planning.
- Stay current with advances in computer vision, multimodal AI, Agentic AI, and Generative AI, and evaluate their applicability to product development.
Required Qualifications
- Bachelor's or Master's degree in Computer Science, Artificial Intelligence, Machine Learning, Electronics, Robotics, or a related discipline.
- 2–4 years of hands-on experience developing machine learning or computer vision applications.
- Robust understanding of machine learning, deep learning, and computer vision fundamentals.
- Hands-on experience training and optimizing deep learning architectures including CNNs, Vision Transformers (ViTs), Transformers, and multimodal models.
- Experience building solutions for object detection, segmentation, tracking, OCR, or 3D perception.
- Experience working with Large Language Models, prompt engineering, Retrieval-Augmented Generation (RAG), AI agents, and tool-calling frameworks.
- Experience fine-tuning or adapting open-source LLMs and Small Language Models.
- Strong programming skills in Python; proficiency in C/C is desirable.
- Hands-on experience with PyTorch and/or TensorFlow.
- Strong understanding of linear algebra, probability, optimization, and machine learning algorithms.
- Good knowledge of camera geometry, calibration, coordinate transformations, projection mathematics, SLAM fundamentals, and point cloud processing.
- Experience with Git, Docker, Linux, CI/CD pipelines, and software engineering best practices.
- Familiarity with cloud platforms such as AWS, Azure, or Google Cloud Platform.
- Excellent analytical, debugging, and problem-solving skills.
📌 Machine Learning Engineer – Multimodal AI & Agentic systems (Bengaluru)
🏢 iMerit Technology
📍 Bengaluru