Client is a pioneering Software Engineering and IT consultancy company, transforming and executing at the intersection of Domain and Technology to create digital leaders for our people, clients, partners, and communities.
About The Job:
- This role is responsible for building reliable, scalable, and governed data pipelines that power analytics, operational reporting, and the Data Intelligence Layer on Databricks.
- The ideal candidate should have strong expertise in Azure Data Engineering, Databricks Lakehouse architecture, streaming data pipelines, and cloud-native data processing frameworks.
Essential Job Functions:
- Build batch pipelines using Delta Lake with optimized table design, partitioning, Z-ORDER, OPTIMIZE, and VACUUM strategies.
- Implement incremental ingestion using Databricks Auto Loader with schema evolution and checkpointing.
- Develop Structured Streaming pipelines with watermarking, state management, and restart safety.
- Implement Lakeflow pipelines for governed and declarative data processing.
- Design replayable and idempotent pipelines supporting protected backfills.
- Optimize Spark workloads using Adaptive Query Execution, shuffle tuning, skew mitigation, and join optimization.
- Build analytics-ready datasets consumed by Databricks SQL, dashboards, Genie, and downstream applications.
- Develop DBSQL models and semantic data layers aligned with Data Intelligence requirements.
- Package and deploy solutions using Databricks Repos and Asset Bundles.
- Support Synapse-to-Databricks migration initiatives and performance benchmarking