Expert-level Python with robust understanding of distributed computing concepts for speed and performance.
Mastery of cloud-based data processing and pipeline orchestration services (AWS, Azure, or equivalent).
Strong understanding of cloud storage solutions across major cloud platforms.
Proficiency with workflow orchestration and scheduling tools (Airflow, Temporal, or equivalent).
Experience automating data pipelines for efficient and reliable end-to-end data processing.
Familiarity with CI/CD principles, Docker, and Kubernetes applied to data engineering.
Strong understanding of data quality concepts — accuracy, consistency, reliability — with hands-on experience implementing validation rules and monitoring systems.