- Design, develop, and maintain scalable ETL/ELT pipelines using Azure Data Factory (ADF) and Databricks.
- Develop efficient data transformation workflows using PySpark.
- Build and optimize batch and incremental data processing pipelines.
- Write efficient SQL queries, stored procedures, and views for data extraction and reporting.
- Develop reusable Python scripts for automation and data processing.
- Integrate data from multiple sources, including databases, APIs, cloud storage, and files.
- Optimize Spark jobs for performance, scalability, and cost efficiency.
- Implement data quality checks, logging, monitoring, and error handling.
- Work with Delta Lake, partitioning, and data optimization techniques.
- Collaborate with business teams to understand data requirements and deliver reliable solutions.
- Troubleshoot production issues and perform root cause analysis.
- Follow coding standards, version control, and CI/CD best practices.
Required SkillsTechnical Skills
- Strong experience with Azure Databricks.
- Hands-on experience with Azure Data Factory (ADF).
- Strong programming skills in Python.
- Solid coding experience in PySpark.
- Advanced SQL skills, including query optimization and performance tuning.
- Experience with Delta Lake and Spark SQL.
- Knowledge of Azure Data Lake Storage (ADLS Gen2).
- Experience with Git or Azure DevOps for source control.
- Understanding of data warehousing concepts and dimensional modeling.
- Experience with REST APIs and data integration.
Preferred Skills
- Experience with Azure Synapse Analytics.
- Knowledge of Microsoft Fabric.
- Experience with streaming technologies such as Spark Structured Streaming or Kafka.
- Exposure to Power BI.
- Familiarity with CI/CD pipelines using Azure DevOps.
- Understanding of Agile/Scrum methodologies.