AI Document Intelligence (Kolkata)

AI Document Intelligence (Kolkata)

11 Sep
|
ARC Document Solutions
|
Kolkata

11 Sep

ARC Document Solutions

Kolkata

Job Title: AI/ML Research Engineer – Document Intelligence (Architectural & MEP Drawings) Location: Kolkata Experience: 2–3 Years Qualification: PhD

Role Overview: We are seeking a research-focused AI/ML professional with a PhD and strong expertise in Computer Vision, Document Intelligence, Multimodal AI, and Machine Learning to work on advanced understanding of architectural, structural, and MEP (Mechanical, Electrical, Plumbing) drawings. The role involves researching and developing intelligent systems capable of extracting and reasoning over text, symbols, dimensions, entities, and spatial relationships within complex engineering drawings. The candidate will contribute to the development of a Small Language Model (SLM) capable of answering questions grounded in architectural and MEP documents by combining document layout understanding, visual recognition, spatial reasoning, and domain-specific language modeling.

This role is particularly suited to a researcher who has demonstrated expertise through a PhD thesis, research publications, prototypes, or applied research in Document AI, Computer Vision, Multimodal AI, Layout Understanding, or related areas.

Key Responsibilities 1. Research in Drawing &

- Document Understanding Research and develop methods for understanding architectural, structural, and MEP drawings. Develop techniques for extracting text, dimensions, symbols, annotations, callouts, legends, and engineering components. Work with both vector PDFs/CAD drawings and raster/scanned documents. Develop spatial and layout-aware representations of complex engineering drawings. Research methods for converting drawings into structured, machine-readable representations. 2.

Computer

Vision &

• Multimodal AI Research and implement computer vision approaches for engineering drawing analysis. Develop models for symbol detection, object recognition, classification, segmentation, and OCR. Investigate multimodal approaches that combine visual information with textual and structural information. Develop spatial grounding techniques linking drawing coordinates with extracted entities and semantic information.

Explore

Vision-Language Models and other multimodal architectures for engineering document understanding. 3.

Small Language Model

Research &





• Development Research, fine-tune, and optimize Small Language Models (SLMs) for domain-specific architectural and MEP Question Answering. Develop methodologies for transforming extracted drawing information into high-quality training and instruction-tuning datasets. Research approaches for domain adaptation, fine-tuning, knowledge grounding, and model distillation. Develop models capable of answering questions involving dimensions, component counts, component locations, relationships, and drawing-specific information.
- Dataset &
- Benchmark Development Develop strategies for collecting, annotating, and validating engineering drawing datasets. Design ground-truth datasets for architectural and MEP drawing understanding. Develop evaluation benchmarks for document extraction, visual grounding, spatial reasoning, and domain specific QA. Conduct experiments and analyze model performance using appropriate statistical and mathematical evaluation Registered methods.
- Research &
- Innovation Conduct research on emerging approaches in Document AI, Layout Understanding, Vision-Language Models, Multimodal AI, SLMs, and spatial reasoning. Review and implement relevant research papers and state-of-the-art techniques. Design experiments to validate current approaches. Document research methodology, experiments, results, and findings for reproducibility and internal knowledge sharing. Contribute to research publications, patents, technical documentation, and intellectual property initiatives where applicable.

Required Skills &

Qualifications Educational Qualification PhD in Computer Science, Artificial Intelligence, Machine Learning, Computer Vision, Document Intelligence, Data Science, or a closely related field.

Professional Experience 2–3 years of relevant industry/research experience in AI/ML, Computer Vision, Document Intelligence, Multimodal AI, NLP,



or related areas. Candidates should have a strong research background demonstrated through a PhD thesis, publications, research projects, prototypes, or production-oriented research work. Relevant experience in document intelligence, engineering drawing understanding, CAD analysis, layout-aware AI, or multimodal systems is highly desirable.

Technical Skills

Strong programming skills in Python. Hands-on experience with PyTorch and/or TensorFlow. Strong understanding of Machine Learning, Deep Learning, Computer Vision, NLP, and Document AI.

Experience with document AI architectures/frameworks such as LayoutLM, Donut, PP-StructureV3, Docling, or equivalent. Strong understanding of OCR, object detection, image segmentation, classification, and visual feature extraction.

Experience with vector/raster PDF processing using tools such as PyMuPDF and OpenCV. Familiarity with CAD formats such as DWG/DXF is an advantage.

Experience with SLM/LLM fine-tuning, model distillation, RAG, domain adaptation, and multimodal learning. Strong mathematics and statistics knowledge is mandatory, including linear algebra, probability, statistics, optimization, mathematical modeling, and mathematical foundations of machine learning and computer vision.

Preferred Skills Research or project experience involving architectural, construction, structural, or MEP drawings. Research publications in Document AI, Computer Vision, Layout Understanding, Multimodal AI, or Vision Language Models.

Experience working with CAD/BIM or engineering drawing datasets.

Experience developing domain-specific QA or multimodal reasoning systems.

Experience creating research datasets and evaluation benchmarks. Strong understanding of current research trends in Vision-Language Models, Document Intelligence, and Small Language Models.

Key Competencies

Exceptional mathematical and analytical ability Strong research and experimentation mindset Strong understanding of AI/ML fundamentals Ability to independently investigate complex technical problems Strong problem-solving and critical-thinking skills Ability to translate research findings into practical AI/ML solutions Strong technical communication and documentation

📌 AI Document Intelligence (Kolkata)
🏢 ARC Document Solutions
📍 Kolkata

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: ai document intelligence (kolkata) / kolkata