02 Oct
|
Geektrust
|
Bengaluru
02 Oct
Geektrust
Bengaluru
Mid Data Engineer – GCP & Spark
- 5–8 years of experience in Data Engineering, Big Data Engineering, or related roles.
Role Overview
We are seeking a highly skilled Senior Data Engineer to design, develop, and support scalable data pipelines and large-scale data processing solutions within the Google Cloud Platform (GCP) ecosystem. The ideal candidate will have strong hands-on experience with Spark, BigQuery, Dataproc, SQL, and cloud-based data engineering practices.
This role requires a strong understanding of end-to-end data pipeline architecture, system design, debugging, performance optimization, and CI/CD practices. The engineer will work closely with data engineering, analytics, and platform teams to build reliable, scalable, and high-performing data solutions.
Key Responsibilities
- Design, develop, and maintain scalable data pipelines using Apache Spark (PySpark or Scala).
- Build and optimize batch and distributed data processing workflows on GCP.
- Work extensively with BigQuery for data transformation, analytics, validation, and reporting.
- Develop and maintain data workflows using GCP services such as Dataproc, BigQuery, and Cloud Storage.
- Design end-to-end data pipeline architectures with a focus on scalability, reliability, and performance.
- Write efficient and optimized SQL queries for large-scale datasets.
- Troubleshoot, debug, and resolve complex data processing and performance issues.
- Implement CI/CD best practices for data engineering workflows.
- Collaborate with cross-functional teams to gather requirements and deliver high-quality data solutions.
- Participate in code reviews, testing, deployment,
and production support activities.
- Monitor pipeline performance and identify opportunities for optimization and continuous improvement.
Required Skills & Qualifications
- Strong hands-on experience with Apache Spark (PySpark and/or Scala).
- Strong SQL skills with experience working on large-scale datasets.
- Hands-on experience with Google Cloud Platform (GCP).
- Practical experience with BigQuery and data warehouse concepts.
- Experience working with Dataproc or similar distributed data processing platforms.
- Solid understanding of data pipelines and end-to-end workflow design.
- Experience designing scalable and reliable data processing solutions.
- Strong debugging, troubleshooting, and problem-solving skills.
- Experience with Git-based development workflows.
- Good understanding of CI/CD practices and deployment processes.
- Experience with Linux/Unix environments and shell scripting.
Preferred Qualifications
- Experience with workflow orchestration tools such as Airflow, Cloud Composer, Luigi, or equivalent.
- Experience with Kafka or other messaging/streaming technologies.
- Knowledge of monitoring, logging, and observability practices.
- Experience with Python for automation and data engineering tasks.
- Exposure to data governance, data quality, and metadata management frameworks.
Soft Skills
- Strong analytical and problem-solving abilities.
- Excellent communication and collaboration skills.
- Ability to work effectively in cross-functional teams.
- Proactive mindset with a focus on operational excellence and continuous improvement.
- Strong ownership and accountability in delivering reliable data solutions.
📌 Mid Data Engineer – GCP & Spark (Bengaluru)
🏢 Geektrust
📍 Bengaluru