23 Aug
|
NTT DATA Business Solutions
|
Hyderabad
23 Aug
NTT DATA Business Solutions
Hyderabad
Data / Retrieval Engineer
Position: Senior Individual Contributor
Experience: 7+ years
Domain: RAG, Enterprise Search, Data Readiness and Retrieval Quality
Openings: 1
Role Overview
We are seeking an experienced Data / Retrieval Engineer to design, build and operate enterprise-grade data ingestion and retrieval capabilities for AI agents and knowledge-search platforms.
The role will focus on transforming enterprise business knowledge into secure, trustworthy and citation-backed context. The engineer will be responsible for improving retrieval quality, grounding, data readiness, latency and operational cost across Retrieval-Augmented Generation and enterprise-search solutions.
Key Responsibilities
- Own the data ingestion and retrieval ecosystem supporting enterprise AI agents.
- Design and build scalable ingestion pipelines for structured, semi-structured and unstructured data.
- Develop indexing solutions using:
- Vector retrieval
- Keyword search
- Hybrid retrieval
- Semantic search
- Implement document chunking, embedding generation, metadata enrichment and indexing strategies.
- Build reranking, grounding and citation-generation mechanisms to improve response accuracy and traceability.
- Convert business documents and knowledge assets into reliable, contextual and reusable data products.
- Implement access-aware retrieval based on users, roles, entitlements and source-system permissions.
- Apply metadata filtering, document lineage, freshness controls and source-level traceability.
- Establish appropriate controls for:
- Personally Identifiable Information
- Sensitive and confidential data
- Data retention
- Security and governance evidence
- Partner with AI-agent, MCP, application and platform engineers to improve retrieval relevance, grounding, latency and infrastructure cost.
- Create retrieval evaluation datasets, including representative queries, expected sources and relevance labels.
- Define and monitor retrieval-quality metrics such as REMAIL_ADDRESS,
PEMAIL_ADDRESS, Mean Reciprocal Rank, NDCG, citation accuracy and groundedness.
- Develop monitoring dashboards, alerts and operational runbooks for retrieval quality, data freshness and model or index drift.
- Troubleshoot production issues related to ingestion failures, missing documents, stale indexes, access-control leakage and poor retrieval relevance.
- Optimise retrieval pipelines for scalability, availability, observability and performance.
Required Skills and Experience
- 7+ years of experience in data engineering, backend engineering, search engineering, machine learning engineering or knowledge-platform development.
- Strong hands-on experience with Python and SQL.
- Production experience implementing Retrieval-Augmented Generation or enterprise-search solutions.
- Solid understanding of:
- Vector databases
- Enterprise-search platforms
- Embedding models
- Chunking strategies
- Keyword and semantic search
- Reranking
- Metadata management
- Experience building data pipelines, APIs and indexing workflows.
- Knowledge of relational databases, document databases or knowledge-retrieval platforms.
- Experience implementing role-based or attribute-based access controls within retrieval systems.
- Strong understanding of data security, privacy, lineage and governance.
- Experience with monitoring, logging, tracing and production observability.
- Ability to work with business, data, AI, security and platform-engineering stakeholders.
- Strong analytical, problem-solving and communication skills.
Preferred Skills
- Experience with large-scale enterprise knowledge bases and multi-source document ingestion.
- Exposure to MCP-enabled applications or agentic AI platforms.
- Experience with document parsing, OCR, table extraction and content normalisation.
- Knowledge of search relevance tuning and learning-to-rank techniques.
- Experience with cloud-based data and AI platforms.
- Familiarity with financial-services data, regulatory content or SP-related business information.
- Experience designing governance evidence, audit trails and data-quality controls.
Indicative Technology Exposure Programming and Data: Python, SQL, APIs, ETL/ELT pipelines
AI and Retrieval: RAG, embeddings, vector search, hybrid search, reranking, grounding, citations
Data Platforms: Vector databases, relational databases, document stores
Search: Enterprise search, keyword search, semantic search, metadata filtering
Governance: Data lineage, access controls, PII management, freshness and retention controls
Operations: Monitoring, logging, observability, quality evaluation and drift detection
Key Success Measures
- Improved retrieval relevance and citation accuracy.
- Reduced hallucination through stronger grounding.
- Reliable enforcement of document and user access permissions.
- Improved freshness and completeness of indexed enterprise data.
- Lower retrieval latency and infrastructure cost.
- Effective identification and remediation of retrieval-quality drift.
- Production-ready monitoring, governance evidence and operational runbooks.
Candidate Red Flags
- Experience limited to prompt engineering without hands-on retrieval implementation.
- No experience evaluating retrieval relevance or grounding quality.
- Limited understanding of data pipelines, metadata or indexing.
- Weak knowledge of security, access control or sensitive-data handling.
- Proof-of-concept experience without production deployment, monitoring or operational ownership.
📌 Data / Retrieval Engineer (Hyderabad)
🏢 NTT DATA Business Solutions
📍 Hyderabad