Key Responsibilities
Design and develop scalable data ingestion, transformation, and processing pipelines using Azure Databricks (PySpark/Scala).
Build and maintain ETL/ELT workflows for structured and unstructured data.
Implement and manage Delta Lake architecture and optimize performance of Spark jobs.
Integrate Azure Databricks with:
Azure Data Factory (ADF)
Azure Data Lake Storage Gen2 (ADLS)
Azure Synapse Analytics
Azure SQL Database
Event Hub/Kafka
Develop reusable frameworks, notebooks, and data engineering best practices.
Optimize Spark workloads through partitioning, caching, indexing, and cluster tuning.
Implement data quality checks, monitoring, and governance standards.
Collaborate with business stakeholders, architects, data scientists, and BI teams.
Support CI/CD implementation using Azure DevOps or GitHub Actions.
Troubleshoot production issues and provide performance improvements.