Python, GCP, BigQuery, ETL pipelines, Data Quality, and Azure
to support an enterprise data platform project.
Key Responsibilities
Develop, maintain, and troubleshoot
ETL/data processing pipelines
Work with
Python, BigQuery, GCS, Dataproc and PySpark
Implement data ingestion, transformation, deduplication and enrichment
Work with
Data Quality (DQ/DQA)
checks and validation frameworks
Handle
SCD Type 2 / historical data versioning
Work with configuration-driven ETL frameworks and metadata-driven processing
Support
Azure ↔ GCP hybrid data pipelines
Work with Azure Storage, Databricks and data ingestion processes
Troubleshoot pipeline, processing,
networking and data-quality issues
Work with
Airflow DAGs
for scheduling and batch processing
Maintain processing status, audit logs and pipeline monitoring
Support migration from artifact-specific ETL scripts toward standardized/configuration-driven pipelines
Required Technical Skills
Must Have:
Python
BigQuery
GCP
GCS
ETL / Data Engineering
Data Quality / Data Validation
SQL
Airflow
Good to Have:
Dataproc / PySpark
Azure
Azure Databricks
FastAPI
Spring Boot
Docker / Kubernetes
Pub/Sub
Service Bus
SCD Type 2
Configuration-driven / metadata-driven ETL