- 4+ years of handson experience in data engineering, data pipelines, ETL/ELT, and large-scale data processing.
- Strong expertise in Google Cloud Platform:
- BigQuery (SQL, performance tuning, partitioning, clustering)
- Dataflow (Apache Beam pipelines)
- Dataproc (Spark, Hadoop, Hive, PySpark)
- Pub/Sub (real-time ingestion)
- Cloud Storage, Cloud Composer (Airflow)
- Cloud SQL / Cloud Spanner
- Robust programming skills in Python, including Pandas and NumPy for data transformations.
- Experience with batch & streaming pipelines.
- Knowledge of data modeling, data lake/lakehouse design, and schema evolution.
- Familiar with workflow orchestration, DAG scheduling, dependency management.
- Experience with CI/CD, Git, unit tests for data pipelines