- Strong hands-on experience with Databricks for building and managing data engineering solutions.
- Expertise in PySpark for large-scale data processing and transformation.
- Advanced SQL skills, including complex queries, joins, aggregations, and performance tuning.
- Proficiency in Python for data processing, automation, and reusable code development.
- Experience developing and optimizing ETL/ELT pipelines for batch and near real-time data workloads.
- Ability to troubleshoot, monitor, and support production data pipelines.
Positive to Have:
- Experience with dbt for data modeling, testing, snapshots, and documentation.
- Knowledge of Apache Airflow for workflow orchestration and scheduling.
- Experience with Azure DevOps, Git, CI/CD pipelines, and deployment automation.
- Familiarity with data warehousing concepts including staging, intermediate, and mart layers.
- Understanding of data lineage, governance, and monitoring best practices.
- Exposure to Structured Streaming and real-time data processing.