- Design, develop, test, deploy and maintain large-scale data pipelines using Airflow on GCP.
- Collaborate with cross-functional teams to gather requirements and design solutions for complex data processing workflows.
- Develop scalable and effective ETL processes using PySpark, SQL, and BigQuery to extract insights from massive datasets.
- Troubleshoot issues related to PubSub messaging queues, Dataproc clusters, and other GCP services.
Job Requirements :
- 4-10 years of experience in Data Engineering with expertise in big data technologies such as Hadoop ecosystem (HDFS), Spark (PySpark) etc.
- Strong understanding of cloud-based platforms like Google Cloud Platform (GCP) including BigQuery, Dataproc, PubSub etc.
- Proficiency in writing complex queries using SQL language for querying large datasets stored in BigQuery.