31 Aug
|
Blute Technologies
|
Bengaluru
31 Aug
Blute Technologies
Bengaluru
GCP Data Engineer – BigQuery, PySpark & Python
Role Summary:
Design, build, and operate large-scale data pipelines and analytics solutions on GCP using SQL, PySpark, and Python.
Partner with architects and business teams to deliver production-grade data products.
Key Responsibilities
Build batch and streaming pipelines on BigQuery, Dataflow, Dataproc, Composer (Airflow), and Pub/Sub.
Develop PySpark transformations on Dataproc / Serverless Spark for large-volume processing.
Write complex, performant SQL for transformation and analytics on multi-billion-row datasets.
Build Python-based ingestion, automation, and reusable frameworks.
Design data models (star/snowflake, SCDs) optimized for BigQuery.
Implement ELT using dbt / Dataform, with DQ, reconciliation, and ABC controls.
Optimize workloads for cost and performance (partitioning, clustering, executor tuning).
Set up CI/CD (Cloud Build / GitHub Actions) and enforce IAM, VPC-SC, CMEK, DLP.
Provide production support — monitoring, RCA, tuning.
Must-Have Skills
Robust SQL — advanced querying, tuning, complex transformations.
PySpark — distributed processing, Spark internals (partitions, shuffles, joins, caching).
Python — hands-on development, automation, framework design, unit testing.
GCP Data Stack — BigQuery, Cloud Storage, Composer/Airflow, Dataproc/Serverless Spark, Dataflow, Pub/Sub.
Data modeling, ELT frameworks (dbt/Dataform), Git, CI/CD.
Positive-to-Have
On-prem → GCP migration experience (Teradata, Oracle, Hadoop).
Banking / financial services background.
Qualifications
B.E./B.Tech/M.Tech in CS, Engineering, or related field.
5–14 years in data engineering, with 3+ years hands-on on GCP.
Work Model
Hybrid — Bangalore location
Pay: ₹1,200,000.00 - ₹2,500,000.00 per year
Work Location: In person
📌 Gcp Data Engineer – Bigquery, Pyspark & Python Bengaluru
🏢 Blute Technologies
📍 Bengaluru