Job Description – GCP Data Engineer
Role Overview
We are looking for an experienced GCP Data Engineer to design, develop, and maintain scalable data pipelines and ETL/ELT frameworks on the Google Cloud Platform. The candidate will work extensively with BigQuery, Python/PySpark, Airflow/Cloud Composer, and Google Cloud Storage to process and transform large volumes of campaign, customer, and clickstream data.
The ideal candidate should have strong hands-on experience in data engineering, excellent SQL and Python skills, and a valuable understanding of cloud-based data pipeline architecture. The role will involve building reliable production pipelines, optimizing BigQuery performance and costs, implementing automation, and supporting critical data workflows.
Key Responsibilities
Data Pipeline Development
- Design, develop, and maintain scalable and reliable data pipelines on GCP .
- Build batch and, where required, near-real-time data ingestion and transformation pipelines.
- Develop robust ETL/ELT frameworks for large volumes of structured and semi-structured data.
- Ingest data from multiple sources into Google Cloud Storage and BigQuery .
- Develop reusable and modular data engineering components.
- Implement appropriate error handling, retry mechanisms, logging, and data validation.
BigQuery & Data Engineering
- Develop complex and optimized SQL queries in BigQuery for data transformation and analysis.
- Optimize BigQuery performance through:
- Partitioning
- Clustering
- Query optimization
- Efficient table design
- Appropriate data types and storage strategies
- Monitor and optimize BigQuery processing costs .
- Design scalable data models suitable for large-scale campaign, customer, and clickstream datasets.
- Troubleshoot data quality, performance, and pipeline-related issues.
Airflow / Cloud Composer
- Develop and maintain DAGs using Apache Airflow / Cloud Composer .
- Implement scheduling, dependency management, retries, failure handling, and alerting.
- Monitor production workflows and proactively resolve failed or delayed jobs.
- Build reusable operators and workflow components where required.
- Ensure critical data pipelines meet agreed SLAs.
Python / PySpark
- Develop data processing and transformation logic using Python and/or PySpark .
- Write clean, reusable, scalable, and production-ready code.
- Optimize PySpark jobs for performance and efficient resource utilization.
- Implement appropriate testing and validation for data transformation processes.
GCP Services
Work extensively with GCP services, particularly:
- Google BigQuery
- Google Cloud Storage (GCS)
- Cloud Composer / Airflow
Exposure to additional GCP services such as Cloud Functions, Pub/Sub, Dataflow, Cloud Run, Secret Manager, IAM, or Cloud Monitoring would be an advantage.
Automation, CI/CD & DevOps
- Automate deployment and execution of data engineering workflows.
- Work with Git-based development and version-control workflows .
- Implement and maintain CI/CD pipelines using tools such as Jenkins or equivalent.
- Follow code review, branching, deployment, and release-management processes.
- Implement monitoring, logging, and alerting for production data pipelines.
Production Support & Troubleshooting
- Monitor daily data pipeline execution and resolve production issues within defined SLAs.
- Investigate pipeline failures, data discrepancies, performance issues, and processing delays.
- Perform root-cause analysis and implement permanent fixes.
- Coordinate with application, analytics, infrastructure, and business teams to resolve data-related issues.
- Participate in production deployments and provide post-deployment support.
Stakeholder Collaboration
- Work closely with Data Analysts, Data Scientists, Product Teams, Business Stakeholders, and Technology Teams to understand data requirements.
- Translate business requirements into scalable technical solutions.
- Communicate technical issues, risks, dependencies, and delivery status effectively.
- Participate in technical discussions, design reviews, and solution development.
Required Technical Skills
Must Have
- 4–8 years of experience in Data Engineering.
- Strong hands-on experience with GCP .
- Strong proficiency in SQL , preferably extensive experience with BigQuery SQL .
- Strong hands-on experience in Python and/or PySpark .
- Experience developing ETL/ELT data pipelines .
- Hands-on experience with BigQuery .
- Hands-on experience with Google Cloud Storage (GCS) .
- Experience with Apache Airflow / Cloud Composer .
- Good understanding of data pipeline architecture and data engineering best practices.
- Experience with Git and CI/CD practices.
- Exposure to Jenkins or similar CI/CD tools .
- Experience in production support, monitoring, troubleshooting, and performance optimization.
Good to Have
- Experience working with campaign, marketing, customer, or clickstream data .
- Experience with large-scale data processing.
- Knowledge of Dataflow / Apache Beam .
- Knowledge of Pub/Sub and event-driven architectures.
- Experience with real-time or streaming data pipelines.
- Knowledge of data warehousing and dimensional data modeling.
- Experience with data quality frameworks and validation.
- Knowledge of GCP IAM and security concepts.
- Experience with Cloud Monitoring / Logging.
- Experience working in Agile/Scrum environments.
Candidate Profile
The ideal candidate should:
- Have strong hands-on technical expertise rather than only theoretical knowledge.
- Be comfortable writing complex SQL and Python/PySpark code .
- Have experience independently designing and developing data pipelines.
- Understand how to build scalable and cost-efficient solutions on GCP .
- Be capable of troubleshooting production issues and taking ownership until resolution.
- Have good analytical and problem-solving skills.
- Be comfortable working with multiple stakeholders and managing delivery timelines.
- Demonstrate good communication and documentation skills.
📌 Process Manager (India)
🏢 eClerx
📍 India