02 Oct
|
V2Soft
|
Bengaluru
- Build the ingestion and chunking pipeline for plan notes, legends, installation details and specifications, and keep the index versioned alongside the vision model releases
- Implement hybrid retrieval combining semantic and keyword search, so exact spec codes and symbol identifiers match literally rather than approximately, with re-ranking over the candidate set
- Enforce citation-or-decline: calibrate the similarity threshold that triggers a decline, and tune it against a regression set to balance false declines against false answers
- Implement answer-faithfulness verification check each claim for support in the cited chunks before the response is returned
- Extract counts, dimensions and spec values verbatim from the source span rather than paraphrasing, and show the snippet beside any number
- Build revision awareness so the bot answers from the current sheet revision and flags superseded ones, plus conflict and discrepancy detection across sheets
- Treat uploaded documents as untrusted data: instruction-like text inside a plan note must never redirect the bot
- Integrate with the HE2 application so AI-generated counts surface with confidence scores and users can accept, reject or reassign them
- Build the curated Q&A; regression suite that scores groundedness and faithfulness on every prompt or model version
What you need
- Strong Python and production experience building RAG systems that are actually in use, not prototypes
- LangChain or equivalent orchestration, and hands-on work with a vector store — Milvus, pgvector, OpenSearch or similar
- Practical grasp of retrieval quality: chunking strategy, hybrid search, re-ranking, and how to measure whether retrieval is the failure or generation is
- Experience with a hosted LLM API; AWS Bedrock and the Claude API specifically is a solid advantage
- FastAPI and async Python service design
- PostgreSQL, and comfort with embedding storage and index lifecycle
- Prompt engineering discipline: evaluation sets, versioned prompts, and regression testing rather than vibes
Nice to have
- Prompt injection and jailbreak mitigation in a production RAG system
- Experience with technical or engineering document corpora
- Cost-aware LLM engineering: prompt caching, batch inference, token budgeting
Building citation and provenance UI alongside the retrieval backend
📌 6 +years only AI Engineer RAG & LLM Systems FastAPI LangChain PyTorc (Bengaluru)
🏢 V2Soft
📍 Bengaluru