We are seeking a skilled Azure Databricks Engineer with solid expertise in PySpark to design, build, and optimize scalable data pipelines on the Azure cloud platform. The ideal candidate will work with large-scale distributed data systems and support advanced analytics and data engineering initiatives.
Key Responsibilities
- Develop and maintain ETL/ELT pipelines using PySpark on Azure Databricks
- Design scalable data solutions using Apache Spark
- Integrate data from multiple sources (Azure Data Lake, SQL DBs, APIs, streaming sources)
- Optimize Spark jobs for performance and cost efficiency
- Work with Azure Data Factory for orchestration
- Implement data transformation, cleansing, and validation logic
- Collaborate with data scientists, analysts, and stakeholders
- Ensure data quality, governance, and security standards
- Monitor and troubleshoot production data pipelines
Required Skills
- Strong experience with PySpark
- Hands-on experience with Azure Databricks
- Good understanding of SQL
- Experience with distributed computing concepts
- Familiarity with Azure Data Lake Storage
- Knowledge of data warehousing concepts
- Experience with version control (Git)
Preferred Skills
- Experience with Delta Lake
- Knowledge of streaming frameworks (Spark Streaming / Kafka)
- Experience with Azure Synapse Analytics
- Python programming beyond PySpark
- CI/CD pipelines (Azure DevOps)
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.