- Robust programming skills in Python and/or Scala, with deep hands-on expertise in PySpark and Spark SQL.
- Solid experience with Delta Lake (ACID transactions, time travel, schema evolution, optimize/vacuum/Z-ordering, liquid clustering).
- Hands-on experience with Lakeflow / Delta Live Tables (DLT) for building declarative, quality-controlled pipelines.
- Working knowledge of Unity Catalog for centralized governance, fine-grained access control, data lineage, and cross-workspace data sharing.
- Experience with Databricks Workflows/Jobs for pipeline orchestration, scheduling, and monitoring.
- Proficiency with at least one major cloud platform (AWS, Azure, or GCP) and its native storage/compute/security services.
- Experience with streaming technologies such as Structured Streaming, Kafka, Event Hubs, or Kinesis.
- Strong SQL skills, including performance tuning, partitioning strategies, and query optimization on large datasets.
- Familiarity with CI/CD for data platforms Databricks Asset Bundles, Git-based version control, Terraform, and automated testing/deployment pipelines.
- Understanding of data modeling concepts (dimensional modeling, medallion/lakehouse architecture) and data warehousing