JOB DESCRIPTION – GCP DATA ENGINEER
ROLE OVERVIEW
We are looking for an experienced GCP Data Engineer to design, develop, and
maintain scalable data pipelines and ETL/ELT frameworks on the Google Cloud
Platform. The candidate will work extensively with BigQuery, Python/PySpark,
Airflow/Cloud Composer, and Google Cloud Storage to process and transform large
volumes of campaign, customer, and clickstream data.
The ideal candidate should have robust hands-on experience in data engineering,
excellent SQL and Python skills, and a good understanding of cloud-based data
pipeline architecture. The role will involve building reliable production
pipelines, optimizing BigQuery performance and costs, implementing automation,
and supporting critical data workflows.
KEY RESPONSIBILITIES
DATA PIPELINE DEVELOPMENT
* Design, develop, and maintain scalable and reliable data pipelines on GCP.
* Build batch and, where required, near-real-time data ingestion and
transformation pipelines.
* Develop robust ETL/ELT frameworks for large volumes of structured and
semi-structured data.
* Ingest data from multiple sources into Google Cloud Storage and BigQuery.
* Develop reusable and modular data engineering components.
* Implement appropriate error handling, retry mechanisms, logging, and data
validation.
BIGQUERY & DATA ENGINEERING
* Develop complex and optimized SQL queries in BigQuery for data transformation
and analysis.
* Optimize BigQuery performance through:
* Partitioning
* Clustering
* Query optimization
* Efficient table design
* Appropriate data types and storage strategies
* Monitor and optimize BigQuery processing costs.
* Design scalable data models suitable for large-scale campaign, customer, and
clickstream datasets.
* Troubleshoot data quality, performance, and pipeline-related issues.
AIRFLOW / CLOUD COMPOSER
* Develop and maintain DAGs using Apache Airflow / Cloud Composer.
* Implement scheduling, dependency management, retries, failure handling, and
alerting.
* M
📌 Process Manager (India)
🏢 eClerx
📍 India