- Design and develop scalable data ingestion, transformation, and processing pipelines using Azure Databricks (PySpark/Scala).
- Build and maintain ETL/ELT workflows for structured and unstructured data.
- Implement and manage Delta Lake architecture and optimize performance of Spark jobs.
- Integrate Azure Databricks with:
- Azure Data Factory (ADF)
- Azure Data Lake Storage Gen2 (ADLS)
- Azure Synapse Analytics
- Azure SQL Database
- Event Hub/Kafka
- Develop reusable frameworks, notebooks, and data engineering best practices.
- Optimize Spark workloads through partitioning, caching, indexing, and cluster tuning.
- Implement data quality checks, monitoring, and governance standards.
- Collaborate with business stakeholders, architects, data scientists, and BI teams.
- Support CI/CD implementation using Azure DevOps or GitHub Actions.
- Troubleshoot production issues and provide performance improvements.