Job Summary
Title: Expert Data Engineer
Shift: 11:00AM - 8:00PM
Work Mode: Hybrid
Location: Bangalore
Introduction to the job:
The AI Data Engineering Lead owns the design, build, governance, and operational readiness of the data pipelines, data products, knowledge assets, and retrieval-ready datasets required to power AI-enabled products.
This role ensures that AI products are not built on fragmented, low-quality, ungoverned, or inaccessible data. The AI Data Engineering Lead works across product, data architecture, AI engineering, platform engineering, security, risk, compliance, MLOps/LLMOps, and operations to ensure that AI products use trusted, permissioned, explainable, and reusable data and knowledge assets.
Some of the things you ll be doing
1. Own AI-ready data engineering
The AI Data Engineering Lead is accountable for engineering data assets that AI products can safely and effectively use.
- Translate AI product needs into data requirements, data pipelines, data products, and knowledge asset requirements.
- Build and manage data pipelines that support AI-enabled applications, RAG solutions, agents, analytics, automation, and decision-support capabilities.
- Ensure data used by AI products is complete, accurate, timely, traceable, and fit for purpose.
- Create reusable data products that can serve multiple products, business units, and AI use cases.
- Partner with product owners and AI engineers to define what data is required for prompts, retrieval, grounding, classification, extraction, recommendations, and workflow automation.
- Ensure AI data assets are engineered for scale, resilience, security, cost efficiency, and production support.
2. Build trusted data products
The role turns raw data into governed, reusable, product-ready data assets.
- Define and build data products for key domains such as client, entity, product, service, transaction, finance, risk, vendor, employee, jurisdiction, and reference data.
- Establish explicit data product ownership, service levels, quality expectations, refresh frequency, and access rules.
- Create data pipelines that support both transactional product needs and AI/analytics needs.
- Define source-of-truth usage and reduce reliance on uncontrolled spreadsheets, local files, and duplicate extracts.
- Ensure data products are documented, discoverable, versioned, and reusable.
- Partner with data stewards to resolve quality, definition, and ownership issues.
3. Engineer data for RAG and knowledge-based AI
The AI Data Engineering Lead ensures documents and knowledge assets can be safely used by AI.
- Prepare policies, procedures, contracts, regulatory content, service playbooks, product documentation, client obligations, and operational knowledge for AI retrieval.
- Define document ingestion, parsing, chunking, embedding, indexing, refresh, and retirement processes.
- Partner with AI engineering to design vector stores, semantic search, retrieval ranking, grounding, and citation patterns.
- Ensure retrieval sources are approved, current, versioned, owned, and access-controlled.
- Validate that AI products retrieve the right content for the right user in the right context.
- Prevent AI products from using obsolete, conflicting, unauthorized, or unapproved knowledge sources.
4. Own data quality and trust controls
AI products amplify data quality issues, so this role establishes trust at the data layer.
- Define data quality rules for critical data elements used by AI products.
- Monitor completeness, accuracy, uniqueness, validity,
consistency, and timeliness.
- Build automated quality checks into pipelines.
- Create exception handling, issue management, and remediation workflows.
- Partner with data owners and product teams to prioritize data quality fixes based on business and AI impact.
- Ensure AI products can identify when data is missing, stale, conflicting, or unreliable.
- Provide evidence of data quality for product acceptance, risk review, and operational readiness.
5. Manage metadata, catalog, and lineage
The AI Data Engineering Lead ensures teams know what data exists, what it means, where it came from, and how it is used.
- Capture metadata for datasets, data products, APIs, pipelines, reports, knowledge assets, vector indexes, and AI retrieval sources.
- Ensure data and knowledge assets are cataloged and discoverable.
- Define business and technical metadata needed for AI use.
- Document lineage from source systems through pipelines, transformations, vector stores, prompts, outputs, and consuming products.
- Support auditability by making source-to-output traceability visible where required.
- Partner with data governance teams to ensure definitions, ownership, sensitivity, and usage rules are documented.
6. Embed data access, privacy, and security controls
The role ensures AI data usage respects permissions, sensitivity, and client obligations.
- Define and implement data access controls in collaboration with data owners and product teams.
- Establish privacy and security controls aligned with regulatory and client obligations.
- Monitor and enforce data access policies across data pipelines and AI systems.
Disclaimer : This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Expert Data Engineer (Bengaluru)
🏢 CSC
📍 Bengaluru