Senior RAG Document AI Engineer (Hyderabad)

Senior RAG Document AI Engineer (Hyderabad)

16 Sep
|
3across
|
Hyderabad

16 Sep

3across

Hyderabad

Senior RAG Document AI Engineer

Job Summary

We are looking for a Senior RAG / Document AI Engineer to design, build, and operate production-grade Python/FastAPI microservices that transform enterprise documents into accurate, cited, and reliable answers.

The role will focus on document processing, intelligent chunking, metadata extraction, hybrid retrieval, reranking, retrieval evaluation, and production operations.

Key Responsibilities

- Design and develop document processing services for PDF, PPTX, DOCX and other enterprise document formats.
- Build pipelines using native text extraction, OCR, and vision-language models.
- Implement table, figure, chart, and infographic extraction.
- Design configurable and versioned chunking strategies based on document type and structure.
- Implement hierarchical, parent-child, and section-aware chunking approaches.
- Develop LLM-based metadata extraction using structured JSON schemas and confidence scoring.
- Build production RAG retrieval services using dense and sparse retrieval.
- Implement BM25, hybrid search, rank fusion, metadata filtering, and cross-encoder reranking.
- Develop context assembly and citation mechanisms for grounded responses.
- Build retrieval evaluation frameworks using Recall@K, nDCG, MRR, groundedness, and golden Q&A; datasets.
- Monitor retrieval quality and use evaluation metrics to drive improvements.
- Design incremental and delta ingestion pipelines and support live corpus re-indexing.
- Build reliable production services with queue/worker patterns, retries, idempotency, and dead-letter handling.
- Implement structured logging, monitoring, tracing, and production support.
- Collaborate with engineering and product teams to improve RAG quality, performance, and reliability.

Required Skills
- 6–10 years of software engineering experience with strong hands-on Python development.
- 2+ years of experience building and operating production Retrieval-Augmented Generation (RAG) or Document AI systems.




- Strong experience with Python and FastAPI in production environments.
- Hands-on experience with hybrid retrieval using dense embeddings and BM25.
- Experience with vector databases and metadata filtering.
- Strong understanding of embedding models, retrieval strategies, rank fusion, and reranking.
- Experience with cross-encoder reranking.
- Practical experience with retrieval evaluation metrics such as Recall@K, MRR, nDCG, and groundedness.
- Robust experience with document processing, OCR, PDF/PPTX/DOCX extraction, and document AI pipelines.
- Experience with LLMs, prompt engineering, structured output, and metadata extraction.
- Hands-on AWS experience with services such as S3, ECS/EKS, Lambda, Bedrock, Textract, or OpenSearch.
- Experience with Docker, pytest, Git, CI/CD, structured logging, and distributed service patterns.
- Strong understanding of asynchronous processing, APIs, queues, workers, retries, and error handling.

Preferred Skills
- Experience working with multimodal RAG or vision-language models.
- Experience in Life Sciences, Pharma, Healthcare, or other regulated environments.
- Knowledge of PII/PHI handling and data security.
- Experience with LangChain or LlamaIndex, with the ability to evaluate when such frameworks are appropriate.
- Experience with MCP or agent-based tool interfaces.
- Experience operating RAG systems at scale with large document collections and high concurrent workloads.

What We Are Looking For
- Strong production engineering mindset.
- Ability to build RAG systems beyond simple API integrations or POCs.
- Strong understanding of retrieval quality and how to measure and improve it.
- Experience handling large-scale document ingestion and re-indexing.
- Ability to troubleshoot and optimize latency, retrieval accuracy, and system reliability.

Keywords RAG, Retrieval Augmented Generation, Generative AI, Document AI, Python, FastAPI, LLM, NLP, OCR, Hybrid Search, BM25, Vector Database, Embeddings, Reranking, Cross Encoder, OpenSearch, AWS, Amazon Bedrock, Textract, Semantic Search, Document Processing, LangChain, LlamaIndex

📌 Senior RAG Document AI Engineer (Hyderabad)
🏢 3across
📍 Hyderabad

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior rag document ai engineer (hyderabad) / hyderabad

Subscribe to this job alert:

Get the latest job offers by email for: senior rag document ai engineer (hyderabad) / hyderabad