Responsibilities:
Design and build conformed dimensions and transactional fact tables for enterprise reporting.
Develop and maintain data pipelines on A zure Databricks using PySpark, Spark SQL, and Delta Lake.
Implement best practices in dimensional modeling, including:
Star schemas
Conformed dimensions
Surrogate and deterministic hash keys (e.g., SHA-256)
Transactional fact design
Translate existing SQL Server and stored procedure logic into effective Spark-based transformations.
Implement incremental data processing and change data capture (CDC) pipelines.
Ensure data reconciliation with source systems to meet accuracy and audit standards.
Collaborate closely with a senior Databricks lead to deliver core data assets within defined timelines.
Required Skills:
8+ years of experience in data engineering, with solid focus on Azure Databricks.
Proficiency in:
PySparkSpark SQLDelta Lake
Deep expertise in dimensional modelling and data warehousing concepts.
Experience with:
Building scalable production pipelinesIncremental data loads and CDC patternsPerformance optimization in distributed systems
Robust understanding of relational databases and ability to interpret complex SQL/stored procedures
Good to have
Experience with:
Lakeflow / Spark Declarative Pipelines or dbt
Unity Catalog
Exposure to financial/General Ledger data models
Familiarity with JD Edwards systems, including:
F0902 (balances)F0911 (journal details)