Role PurposeBuild and operate the data pipelines that feed the Entity Hub. This role lands all six in-scope sources into Fabric, implements standardization and transformation logic, and maintains the data quality checks and monitoring that the entity resolution engine depends on. Reliable, observable ingestion is the foundation the entire programme rests on.Key ResponsibilitiesIngestion development —
build and maintain pipelines to land the six in-scope sources (Secretary of State, D&B;, ARROW, E1, hCue, DocCentral) into the Fabric Bronze/raw layer.Mirroring & CDC —
implement Fabric Mirroring for supported structured sources and establish change-data-capture patterns; implement watermark/incremental load logic where mirroring is unavailable.Raw layer management —
maintain one Delta table per source on an append-only basis, retaining evidence records and full source provenance.Standardization & transformation —
implement name normalization, address parsing and attribute standardization logic in Spark notebooks; support identifier-spine construction.Data quality —
implement data quality checks, validation rules, threshold alerts and exception handling; support reconciliation against source.Pipeline operations —
schedule, monitor and troubleshoot pipeline runs; investigate failures and performance issues; maintain run documentation.Performance tuning —
optimise Spark jobs, Delta file sizes, partitioning and pipeline efficiency to manage Fabric capacity consumption.Documentation —
produce and maintain source-to-target mappings, transformation logic documentation and lineage records.Required Skills & ExperienceSkill AreaSpecific RequirementsCore EngineeringPython, PySpark, advanced SQL, Delta Lake, distributed data processingMicrosoft FabricData Factory pipelines and Copy Activity, Lakehouse, OneLake, Spark notebooks,
Environments, Mirroring, ShortcutsData IntegrationBatch and incremental ingestion, CDC patterns, watermarking, reprocessing strategies, schema-on-read for wide-ranging formatsData QualityValidation rule implementation, completeness/accuracy checks, alerting, exception workflows, reconciliationModellingBronze/Silver/Gold medallion layering, cleansing and conformance, standardization of names, addresses, dates and codesOps & GovernancePipeline monitoring, lineage and metadata capture, access controls, technical documentationMust-Have Qualifications4+ years hands-on data engineering with strong PySpark and SQLProduction experience building ingestion pipelines from multiple heterogeneous sourcesWorking knowledge of Delta Lake and medallion/lakehouse architectureExperience implementing incremental loads and CDC-style processingExperience implementing data quality checks and troubleshooting pipeline failuresNice-to-HaveMicrosoft Fabric hands-on experience (Mirroring, Copy Jobs, Environments)Exposure to entity/master data standardization (name and address parsing)Familiarity with libraries such as Great Expectations for data qualityExperience optimising for Fabric capacity/CU consumptionKey Deliverables OwnedOperational ingestion pipelines for all agreed sourcesBronze/raw layer with one Delta table per source and CDC retainedStandardization and parsing transformation logicData quality checks, monitoring and exception handlingSource-to-target mapping and run documentationDual Role / Complementary SkillsComplementary with the Entity Resolution engineering workstream — both are PySpark-on-Fabric disciplines, so this role can cross-train on Splink tuning and candidate-pair generation to provide cover. Also supports the Sr. Data Engineer (Lead) on identifier-spine construction, and can assist the VectorDB Engineer with document/attribute preparation in Phase 2.
📌 Data Engineer Fabric (Mumbai)
🏢 EXL
📍 Mumbai