- Design, develop, and maintain scalable data pipelines using Azure Databricks.
- Develop ETL/ELT processes for ingesting, transforming, and loading large-scale datasets.
- Build and optimize Spark-based applications using PySpark, Scala, or SQL.
- Integrate data from multiple sources, including databases, APIs, streaming platforms, and cloud storage.
- Develop and manage Delta Lake architectures for reliable and high-performance data processing.
- Implement data quality checks, monitoring, and performance optimization techniques.
- Collaborate with business analysts, data architects, and stakeholders to understand data requirements.
- Design and implement data models for analytics and reporting.
- Ensure data security, governance, and compliance standards are maintained.
- Participate in code reviews, testing, deployment, and production support activities.
- Troubleshoot and resolve performance bottlenecks within Azure Databricks environments.
- Develop reusable frameworks and best practices for data engineering solutions.