Role & responsibilities
Design, develop, and maintain scalable data pipelines using Azure Databricks.
Develop ETL/ELT processes for ingesting, transforming, and loading large-scale datasets.
Build and optimize Spark-based applications using PySpark, Scala, or SQL.
Integrate data from multiple sources, including databases, APIs, streaming platforms, and cloud storage.
Develop and manage Delta Lake architectures for reliable and high-performance data processing.
Implement data quality checks, monitoring, and performance optimization techniques.
Collaborate with business analysts, data architects, and stakeholders to understand data requirements.
Design and implement data models for analytics and reporting.
Ensure data security, governance, and compliance standards are maintained.
Participate in code reviews, testing, deployment, and production support activities.
Troubleshoot and resolve performance bottlenecks within Azure Databricks settings.
Develop reusable frameworks and best practices for data engineering solutions.