29 Aug
|
Blute Technologies
|
Bengaluru
29 Aug
Blute Technologies
Bengaluru
GCP Data Engineer – BigQuery, PySpark & Python
Role Summary
Design, build, and operate large-scale data pipelines and analytics solutions on GCP using SQL, PySpark, and Python.
Partner with architects and business teams to deliver production-grade data products.
Key Responsibilities
- Build batch and streaming pipelines on BigQuery, Dataflow, Dataproc, Composer (Airflow), and Pub/Sub.
- Develop PySpark transformations on Dataproc / Serverless Spark for large-volume processing.
- Write complex, performant SQL for transformation and analytics on multi-billion-row datasets.
- Build Python-based ingestion, automation, and reusable frameworks.
- Design data models (star/snowflake, SCDs) optimized for BigQuery.
- Implement ELT using dbt / Dataform, with DQ, reconciliation, and ABC controls.
- Optimize workloads for cost and performance (partitioning, clustering, executor tuning).
- Set up CI/CD (Cloud Build / GitHub Actions) and enforce IAM, VPC-SC, CMEK, DLP.
- Provide production support — monitoring, RCA, tuning.
Must-Have Skills
- Strong SQL — advanced querying, tuning, complex transformations.
- PySpark — distributed processing, Spark internals (partitions, shuffles, joins, caching).
- Python — hands-on development, automation, framework design, unit testing.
- GCP Data Stack — BigQuery, Cloud Storage, Composer/Airflow, Dataproc/Serverless Spark, Dataflow, Pub/Sub.
- Data modeling, ELT frameworks (dbt/Dataform), Git, CI/CD.
Positive-to-Have
- On-prem → GCP migration experience (Teradata, Oracle, Hadoop).
- Banking / financial services background.
Qualifications
- B.E./B.Tech/M.Tech in CS, Engineering, or related field.
- 5–14 years in data engineering, with 3+ years hands-on on GCP.
Work Model
Hybrid — Bangalore location
Pay: ₹1,200,000.00 - ₹2,500,000.00 per year
Work Location: In person
📌 GCP Data Engineer – BigQuery, PySpark & Python (Bengaluru)
🏢 Blute Technologies
📍 Bengaluru