03 Sep
|
MindBrain
|
Gurugram
03 Sep
MindBrain
Gurugram
Senior Data Engineer – GCP
Experience: 5+ Years
Role: Senior Data Engineer
Domain: Cloud Data Engineering
Platform: Google Cloud Platform (GCP)
Role Summary
We are seeking a Senior Data Engineer with 5+ years of experience in designing, developing, and maintaining scalable, cloud-based data solutions. The ideal candidate will have strong hands-on expertise in SQL, Python, PySpark, and Google Cloud Platform (GCP).
The candidate will be responsible for building enterprise-grade ETL/ELT pipelines, implementing scalable batch and real-time data solutions, optimizing data processing workloads, and ensuring high standards of data quality, reliability, and performance.
You will collaborate closely with Product Owners, Engineering teams, Business SMEs, and other stakeholders to understand requirements and deliver robust, scalable, and high-quality data products.
Mandatory Skills
- Python – Advanced
- SQL – Advanced
- PySpark
- Google Cloud Platform (GCP)
- BigQuery
- Cloud Composer / Apache Airflow
- Dataproc
- Pub/Sub
- Google Cloud Storage (GCS)
- ETL / ELT Development
- Batch & Streaming Data Processing
- Data Pipeline Development & Optimization
- Git / GitHub
- CI/CD Concepts
Key Responsibilities
- Design, develop, and maintain scalable ETL/ELT data pipelines using Python and PySpark.
- Build and optimize cloud-native data solutions on Google Cloud Platform.
- Develop optimized SQL queries, transformation logic, and data models in BigQuery.
- Design and maintain workflow orchestration using Cloud Composer / Apache Airflow.
- Develop, optimize, and troubleshoot Spark applications running on Dataproc.
- Design and implement event-driven data ingestion pipelines using Google Pub/Sub.
- Build and manage both batch and real-time data processing pipelines.
- Develop data solutions for Data Warehouses and Data Lakes.
- Implement data validation, reconciliation, exception handling, monitoring,
and data quality controls.
- Optimize pipeline performance, query execution, resource utilization, and processing costs.
- Develop reusable data engineering frameworks, utilities, and components to improve engineering productivity.
- Implement logging, monitoring, alerting, and troubleshooting mechanisms for production data pipelines.
- Participate in code reviews and follow established engineering, coding, and development best practices.
- Create and maintain comprehensive technical documentation for data pipelines, workflows, data models, and architecture.
- Collaborate with Product Owners, Business SMEs, Engineering teams, and other stakeholders to translate business requirements into scalable technical solutions.
Google Cloud & Data Engineering Expertise
Hands-on experience with Google's data ecosystem, including:
- BigQuery
- Dataproc
- Dataflow
- Pub/Sub
- BigTable
- Cloud Spanner
- Cloud SQL
- AlloyDB
- Google Cloud Storage
- Cloud Composer / Apache Airflow
- Cloud Scheduler
Experience designing and implementing batch and real-time data pipelines, data migration solutions, and scalable data-layer architectures across the GCP ecosystem is highly desirable.
Data Warehouse & Data Modeling
Strong understanding of:
- Data Warehouse Architecture
- Data Lake vs. Data Warehouse
- Star Schema
- Snowflake Schema
- Fact & Dimension Modeling
- Slowly Changing Dimensions (SCD)
- Partitioning & Clustering
- Data Layer Design
- Data Transformation & Aggregation
- Data Quality & Governance
Big Data & Open-Source Technologies
Experience with one or more of the following:
- Apache Spark
- PySpark / Python
- Spark / Scala
- Apache Hadoop
- Apache Beam
- Apache Airflow
- dbt
DevOps & Engineering Practices
- Git / GitHub
- CI/CD
- Docker
- Version Control
- Code Reviews
- Automated Testing
- Logging & Monitoring
- Production Support & Troubleshooting
Preferred Qualifications
- Experience working on enterprise-scale cloud migration initiatives.
- Experience developing reusable data engineering frameworks and utilities.
- Strong understanding of data governance and data quality practices.
- Exposure to AI-assisted development tools, such as GitHub Copilot.
- Experience with both batch and streaming data architectures.
- Strong analytical, problem-solving, and troubleshooting skills.
- Excellent communication and stakeholder management skills.
- Ability to work effectively with cross-functional Product, Engineering, and Business teams.
Experience Requirements
- 5+ years of professional experience in Data Engineering.
- Strong hands-on experience with SQL, Python, and PySpark.
- Proven experience developing production-grade data pipelines.
- Hands-on experience with Google Cloud Platform (GCP) and its data engineering services.
- Experience designing scalable solutions for Data Warehouses, Data Lakes, batch processing, and real-time streaming.
Ideal Candidate
The ideal candidate is a hands-on Senior Data Engineer who can independently design and deliver scalable data solutions using Python, SQL, PySpark, and GCP, with solid expertise in BigQuery, Dataproc, Cloud Composer/Airflow, Pub/Sub, and Cloud Storage.
The candidate should combine strong technical expertise with a solid understanding of data architecture, data modeling, data quality, cloud engineering, and production operations.
📌 Senior Data Engineer GCP (5+ Yrs) (Gurugram)
🏢 MindBrain
📍 Gurugram