As a Data Engineer, you will help build the data foundation for our agentic AI platform. You will work with senior data architects, AI/ML engineers, and platform engineers to implement data ingestion, transformation, profiling, enrichment, validation, and preparation pipelines across structured and unstructured enterprise data sources.
This is a hands-on engineering role for someone who enjoys working with real-world enterprise data, building reliable pipelines, writing robust Python and SQL, and helping convert raw enterprise information into AI-ready data assets.
Key Responsibilities
Build and maintain data ingestion pipelines for structured enterprise systems such as ERP, CRM, billing, finance, HR, OSS/BSS, ServiceNow, Salesforce, SAP, Oracle, databases, and APIs.
Build pipelines for unstructured and semi-structured data sources such as documents, emails, logs, transcripts, PDFs, spreadsheets, and media metadata.
Develop ETL/ELT workflows using Python, SQL, PySpark, Apache Spark, Airflow, dbt, Dagster, cloud-native services, or equivalent technologies.
Support data profiling routines to identify missing values, duplicates, inconsistent formats, incomplete master data, schema changes, and conflicting records.
Implement data quality checks using frameworks such as Excellent Expectations, dbt tests, AWS Glue DataBrew, custom validation scripts, or equivalent tools.
Support data labelling, contextualization, harmonization, enrichment, and classification workflows required for AI agent configuration.
Prepare data outputs for downstream AI consumption, including embeddings, metadata, semantic tags, graph-ready datasets, and retrieval-ready document chunks.
Work with vector databases and search indexes to support semantic retrieval and hybrid search use cases.
Support knowledge graph and ontology implementation by preparing entities, relationships, attributes, mappings, and validation-ready datasets.
Implement data masking, anonymization, access control, PII handling, and
📌 Data Engineer (Noida)
🏢 Infosys
📍 Noida