Strong hands-on experience in Databricks and PySpark.
Good understanding of Spark architecture, transformations, partitioning, joins, and performance tuning and transformations like SCD1/SCD2.
Solid SQL and data engineering fundamentals.
Experience working with cloud-based data platforms, preferably Azure.
Good problem-solving, troubleshooting, and communication skills.
We are looking for a Data Engineer with strong hands-on experience in Databricks and PySpark to develop and maintain scalable data pipelines and data processing solutions along with transformations like SCD1/SCD2
Key Responsibilities
Develop, enhance, and optimize ETL/ELT pipelines using Databricks and PySpark.
Work with large datasets and implement Spark performance optimization techniques.
Develop complex SQL queries for data transformation and analysis and transformations like SCD1/SCD2.
Work with ADLS and ADF for data storage and pipeline orchestration.
Support and modernize existing Hadoop/Hive data processing workloads.
Develop Shell scripts for automation and operational activities.
Implement and maintain CI/CD pipelines for code deployment and data engineering workflows.
Troubleshoot production issues and ensure pipeline performance, reliability, and data quality.
Collaborate with cross-functional teams to deliver scalable and reliable data solutions.