Job Description:
- Design and build conformed dimensions and transactional fact tables for enterprise reporting.
- Develop and maintain data pipelines on Azure Databricks using PySpark, Spark SQL, and Delta Lake.
- Implement best practices in dimensional modeling, including:
- Star schemas
- Conformed dimensions
- Surrogate and deterministic hash keys (e.g., SHA-256)
- Transactional fact design
- Translate existing SQL Server and stored procedure logic into effective Spark-based transformations.
- Implement incremental data processing and change data capture (CDC) pipelines.
- Ensure data reconciliation with source systems to meet accuracy and audit standards.
- Collaborate closely with a senior Databricks lead to deliver core data assets within defined timelines.