Job Summary
The AI Data Engineering Lead owns the design, build, governance, and operational readiness of the data pipelines, data products, knowledge assets, and retrieval-ready datasets required to power AI-enabled products.
This role ensures that AI products are not built on fragmented, low-quality, ungoverned, or inaccessible data. The AI Data Engineering Lead works across product, data architecture, AI engineering, platform engineering, security, risk, compliance, MLOps/LLMOps, and operations to ensure that AI products use trusted, permissioned, explainable, and reusable data and knowledge assets.
Job Details
Title: Expert Data Engineer
Shift: 11:00AM - 8:00PM
Work Mode: Hybrid
Location: Bangalore
Responsibilities
- Own AI-ready data engineering. The AI Data Engineering Lead is accountable for engineering data assets that AI products can safely and effectively use.
- Build and manage data pipelines that support AI-enabled applications, RAG solutions, agents, analytics, automation, and decision-support capabilities.
- Ensure data used by AI products is complete, accurate, timely, traceable, and fit for purpose.
- Create reusable data products that can serve multiple products, business units, and AI use cases.
- Partner with product owners and AI engineers to define what data is required for prompts, retrieval, grounding, classification, extraction, recommendations, and workflow automation.
- Ensure AI data assets are engineered for scale, resilience, security, cost efficiency, and production support.
- Define data products for key domains such as client, entity, product, service, transaction, finance, risk, vendor, employee, jurisdiction, and reference data.
- Establish explicit data product ownership, service levels,
quality expectations, refresh frequency, and access rules.
- Create data pipelines that support both transactional product needs and AI/analytics needs.
- Define source-of-truth usage and reduce reliance on uncontrolled spreadsheets, local files, and duplicate extracts.
- Ensure data products are documented, discoverable, versioned, and reusable.
- Partner with data stewards to resolve quality, definition, and ownership issues.
- Prepare policies, procedures, contracts, regulatory content, service playbooks, product documentation, client obligations, and operational knowledge for AI retrieval.
- Define document ingestion, parsing, chunking, embedding, indexing, refresh, and retirement processes.
- Partner with AI engineering to design vector stores, semantic search, retrieval ranking, grounding, and citation patterns.
- Ensure retrieval sources are approved, current, versioned, owned, and access-controlled.
- Validate that AI products retrieve the right content for the right user in the right context.
- Prevent AI products from using obsolete, conflicting, unauthorized, or unapproved knowledge sources.
- Define data quality rules for critical data elements used by AI products.
- Monitor completeness, accuracy, uniqueness, validity, consistency, and timeliness.
- Build automated quality checks into pipelines.
- Create exception handling, issue management,
and remediation workflows.
- Partner with data owners and product teams to prioritize data quality fixes based on business and AI impact.
- Ensure AI products can identify when data is missing, stale, conflicting, or unreliable.
- Provide evidence of data quality for product acceptance, risk review, and operational readiness.
- Capture metadata for datasets, data products, APIs, pipelines, reports, knowledge assets, vector indexes, and AI retrieval sources.
- Ensure data and knowledge assets are cataloged and discoverable.
- Define business and technical metadata needed for AI use.
- Document lineage from source systems through pipelines, transformations, vector stores, prompts, outputs, and consuming products.
- Support auditability by making source-to-output traceability visible where required.
- Partner with data governance teams to ensure definitions, ownership, sensitivity, and usage rules are documented.
- Implement role-based, attribute-based, jurisdictional, client-specific, and purpose-based access controls.
- Ensure AI products only use data that users, applications, models, and agents are authorized to access.
- Partner with security and privacy teams to classify data sensitivity.
- Prevent sensitive data from being exposed through prompts, logs, embeddings, outputs, or retrieved content.
- Ensure data minimization, retention, masking, encryption, and audit logging requirements are met.
- Design data access patterns that work across applications, APIs, data products, vector stores, and AI agents.
- Support security and privacy reviews for AI-enabled products.
- Monitor pipeline reliability, latency, and freshness.
📌 Expert Data Engineer (Bengaluru)
🏢 CSC
📍 Bengaluru