24 Sep
|
Shapoorji Pallonji Finance Private
|
Mumbai
24 Sep
Shapoorji Pallonji Finance Private
Mumbai
Shapoorji Pallonji Group
Experience: 5+ years
Location: Mumbai
About the Role
We build enterprise intelligence platforms that extract and reason over information contained within documents, images, diagrams, charts, drawings, video, and other visual media.
We are looking for a Senior Computer Vision Engineer specialising in Multimodal AI and Document Intelligence to design, build, and deploy production-grade systems that make complex visual information machine-understandable.
The role covers document intelligence, OCR, diagram and chart interpretation, vision-language models, multimodal retrieval, image restoration, and video understanding. You will work across the complete lifecycle, including data preparation, model selection, training, fine-tuning, evaluation, optimisation, deployment, and production monitoring.
This is a hands-on product engineering role, not a research-only position. The person must have experience deploying computer vision or multimodal machine learning solutions into production environments used by enterprise customers.
Key Responsibilities
- Document Intelligence and OCR
- Build systems for document layout analysis, region detection, reading-order reconstruction, table extraction, and structural understanding across scanned and digitally generated documents.
- Develop and improve multilingual OCR pipelines, including Indic-language documents, handwriting, stamps, seals, signatures, watermarks, redactions, and annotations.
- Improve OCR quality through engine selection, preprocessing, model ensembling, post-correction, and confidence calibration.
2. Visual and Diagram Understanding
- Build solutions to interpret diagrams, flowcharts, organisation charts, network diagrams, schematics, technical drawings, maps, timelines, and scientific figures.
- Develop techniques for symbol recognition, connector tracing, label association, dimension extraction, legend interpretation, and relationship identification.
- Build visual comparison capabilities to identify meaningful changes across document, drawing, or diagram versions.
3. Charts, Tables and Data Visualisation
- Extract structured information from charts, dashboards, infographics, and complex tables.
- Develop methods for axis detection, scale inference, legend association, series separation, value extraction, and borderless table reconstruction.
- Implement confidence scoring where the source image does not allow precise extraction.
4. Multimodal AI and Visual Retrieval
- Build vision-language model applications for grounded question answering across documents, images, diagrams, and video frames.
- Design multimodal retrieval pipelines using visual embeddings, hybrid retrieval, reranking, and late-interaction approaches such as ColPali or ColQwen.
- Fine-tune open-weight vision-language models using LoRA, QLoRA, instruction tuning, or other parameter-efficient techniques.
- Build multimodal RAG and GraphRAG solutions that preserve source context, structure, relationships, and provenance.
5. Image Processing and Synthetic Data
- Build image preprocessing and restoration pipelines covering de-skewing, de-noising, de-warping, shadow removal, moiré reduction, binarisation, and super-resolution.
- Use generative methods, including diffusion models and GANs, for restoration, inpainting, augmentation, and synthetic data creation.
- Generate degraded and long-tail visual samples to test model robustness and improve performance where labelled training data is limited.
6. Video and Temporal Understanding
- Develop capabilities for keyframe extraction, scene segmentation, object tracking, text extraction, and temporal event detection.
- Extract slides, screens, documents, and other relevant visual information from recorded video and screen captures.
- Build solutions for video summarisation, action recognition, and event-based retrieval.
7. Evaluation and Production Deployment
- Define evaluation metrics before model development and build golden datasets, regression tests, hallucination checks, structural fidelity measures, and field-level accuracy benchmarks.
- Conduct systematic failure analysis across resolution, language, layout, skew, visual density, and domain variation.
- Implement confidence scoring and human review workflows for uncertain outputs.
- Optimise models for accuracy, latency, throughput, memory consumption, infrastructure cost, and production reliability.
- Deploy and monitor models in cloud, on-premises, or restricted client environments.
Required Experience and Skills
- 5+ years of relevant experience in Computer Vision, Machine Learning, Document AI, Multimodal AI, or Applied AI Engineering.
- Strong Python programming skills and deep hands-on experience with PyTorch.
- Proven experience building and deploying production-grade computer vision or multimodal machine learning systems.
- Hands-on experience with document AI, OCR, layout analysis, image processing, or visual information extraction.
- Strong working knowledge of:
- Classical computer vision, including geometric transformations, morphology, contour analysis, feature matching, and image registration
- CNN architectures such as ResNet, EfficientNet, and U-Net
- Detection and segmentation models such as YOLO, DETR, Mask R-CNN, and SAM
- Vision transformers and vision-language models such as ViT, CLIP, Qwen-VL, InternVL, Donut, or Pix2Struct
- Generative models including diffusion models, GANs,
and VAEs
- Experience training, fine-tuning, evaluating, and deploying models using GPU infrastructure.
- Strong understanding of model throughput, batching, memory optimisation, inference cost, and production scalability.
- Ability to evaluate when classical computer vision, a specialised model, or a large vision-language model is the most effective solution.
- Strong analytical ability with a disciplined approach to experimentation, benchmarking, and measurable improvement.
Preferred Experience
- Multimodal RAG, GraphRAG, knowledge graphs, or Neo4j-based retrieval.
- Visual retrieval using multimodal embeddings, hybrid search, reranking, or late-interaction models.
- Diagram, schematic, engineering drawing, CAD, BIM, medical imaging, scientific imaging, or geospatial understanding.
- Indic-language or multilingual document processing.
- Video analytics, temporal understanding, or screen-recording intelligence.
- Model serving and optimisation using ONNX, TensorRT, vLLM, quantisation, or inference batching.
- Experience deploying solutions within on-premises, private-cloud, or air-gapped environments.
- Data labelling strategy, annotation workflows, active learning, and synthetic dataset generation.
- Open-source contributions, publications, patents, or demonstrable technical work in computer vision, document AI, or multimodal systems.
What We Are Looking For
- A hands-on engineer who has taken models beyond experimentation and deployed them in production.
- Strong problem-solving ability across unfamiliar document, image, diagram, and video formats.
- A quality-first mindset supported by measurable evaluation and structured failure analysis.
- Ability to balance model accuracy with latency, infrastructure cost, security, and scalability.
- Strong ownership across model development, deployment, monitoring, and continuous improvement.
- Ability to collaborate effectively with Product, Data Science, Engineering, and enterprise implementation teams.
Success in This Role
Success will be measured by:
- Accuracy and structural fidelity of extracted visual information.
- Effectiveness of retrieval and question answering across visual content.
- Improvement against defined evaluation benchmarks and golden datasets.
- Reliability, cost efficiency, and scalability of deployed models.
- Reduction in manual review through well-calibrated confidence scoring.
- Ability to bring new visual formats into production-grade processing pipelines.
- Business adoption and performance of visual intelligence capabilities in enterprise products.
Significant Screening Requirement
Candidates should have experience delivering production-grade computer vision, document AI, or multimodal AI solutions. Experience limited to academic research, notebooks, prototypes, hackathons, or proofs of concept will not be sufficient.
📌 Senior Computer Vision Engineer - Multimodal AI (Mumbai)
🏢 Shapoorji Pallonji Finance Private
📍 Mumbai