- Hands-on experience with Databricks platform, including:
- PySpark / Spark SQL for data processing and transformation
- Delta Lake for ACID-compliant data storage
- Notebooks for workflow orchestration and collaborative development
- Robust programming skills in Python and SQL.
- Experience in data ingestion from batch and streaming sources.
- Knowledge of ETL/ELT design patterns and data pipeline optimization.
- Familiarity with cloud data storage and compute environments (AWS, Azure, or GCP).
- Basic understanding of workflow orchestration tools (Airflow, Databricks Jobs).
- Exposure to DevOps concepts, CI/CD pipelines, and version control (Git).
- Awareness of data security, access control, and compliance considerations.
- Understanding of performance tuning, partitioning, and caching strategies in Spa