We are looking for an experienced Azure Databricks Data Engineer to design, develop, and maintain scalable data engineering solutions using Databricks, PySpark, Spark SQL, and Python. The role involves building reliable data pipelines, transforming large datasets, and supporting enterprise data platforms on Azure.
Key Responsibilities
- Design, develop, and maintain scalable data pipelines using Azure Databricks.
- Develop data transformation and processing workflows using PySpark and Spark SQL.
- Write efficient and reusable Python code for data engineering solutions.
- Work with large and complex datasets to perform data extraction, transformation, and loading.
- Develop and optimize Spark jobs for performance, scalability, and reliability.
- Implement data quality checks, validation, and error-handling mechanisms.
- Troubleshoot and optimize Databricks notebooks, jobs, and Spark workloads.
- Collaborate with data architects, analysts, and application teams to understand data requirements.
- Follow best practices for data engineering, coding standards, security, and documentation.
- Support production deployments and resolve data pipeline issues.
Mandatory Skills
- Strong hands-on experience with Azure Databricks
- Strong PySpark development experience
- Strong knowledge of Spark SQL
- Solid programming experience in Python
- Good understanding of data engineering concepts and ETL/ELT pipelines
- Experience working with large-scale data processing using Apache Spark
- Solid SQL and data transformation skills
- Experience in performance tuning and optimization of Spark/Databricks workloads
Good to Have
- Experience with Azure Data Factory
- Knowledge of Azure Data Lake Storage (ADLS)
- Experience with Delta Lake and Delta tables
- Knowledge of CI/CD and Databricks deployment practices