25 Sep
|
Tiger Analytics
|
India
25 Sep
Tiger Analytics
India
Design, build, and maintain robust cloud data pipelines using Azure Data Factory (ADF), Azure Databricks, PySpark, and Spark SQL.
Implement and manage the Medallion Architecture — moving and transforming data through Bronze (raw/audit), Silver (cleansing/dedup/SCD), and Gold (business aggregates) layers.
Perform complex data transformations, cleansing, deduplication, and incremental loads using Delta MERGE, supporting both Slowly Changing Dimensions (SCD Type 1 and Type 2).
Optimize Spark workloads by tuning shuffle partitions, managing memory to prevent OOM errors, and leveraging Adaptive Query Execution (AQE).
Apply optimized join strategies, including broadcast joins for small datasets and salting techniques to handle data skew.
Implement robust exception handling,
file dependency validation, and asynchronous batch processing to ensure pipeline reliability.
Ensure high data quality through schema enforcement, schema evolution handling, and validation against expected target criteria.
Monitor, troubleshoot, and resolve production job failures by analysing cluster scaling behaviour, Spark UI metrics, and physical query plans.
Build and maintain CI/CD pipelines for Databricks using Git, Azure DevOps, and Databricks Asset Bundles (DABs), with setting-specific parameterization for Dev, Test, and Prod.
📌 Databricks Azure Dataengineer & Sr Data Engineer Pune (India)
🏢 Tiger Analytics
📍 India