21 Aug
|
RandomTrees
|
Secunderabad
21 Aug
RandomTrees
Secunderabad
Role & Responsibilities:
Design, develop, and maintain ELT/ETL pipelines on GCP using Dataflow/Beam, Dataproc/Spark, and Airflow/Composer.
Model and optimize datasets in BigQuery using partitioning, clustering, materialized views, and UDFs.
Build streaming and near-real-time data ingestion using Pub/Sub, Dataflow, and CDC where applicable.
Implement data quality checks, validation frameworks, and SLAs, monitoring pipelines via Cloud Monitoring/Logging.
Optimize performance and cost across GCS, Dataproc autoscaling, and BigQuery slot usage.
Contribute to coding, CI/CD standards, observability, documentation, and perform code reviews.
Partner with Analytics, BI, and ML teams to productize datasets and ensure solid data contracts.
Support production operations, including on-call rotations for critical pipelines.
Preferred Candidate Profile:
Hands-on expertise in GCP data stack: BigQuery, Dataflow (Apache Beam), Dataproc, Cloud Storage, Pub/Sub, Cloud Composer (Airflow).
Strong experience with Spark (PySpark or Scala) for batch processing.
Solid Airflow DAG design knowledge (idempotent tasks, backfills, retries, SLAs).
Advanced SQL and data modeling (star/snowflake schemas, slowly changing dimensions, partition strategies).
Proficiency in Python (preferred) or Scala/Java.
Experience with Git and CI/CD tools (Cloud Build, GitHub Actions, GitLab CI).
Familiarity with GCP security and governance (IAM, service accounts, secrets management, VPC-SC).
Solid debugging, problem-solving, and communication skills.
Positive-to-Have:
Snowflake (migration, performance tuning, tasks/streams).
Terraform for infrastructure as code on GCP.
Kafka or other streaming tools; Cloud Run/Functions for glue services.
BI exposure (Looker, Tableau, Power BI).
GCP Professional Data Engineer certification.
📌 Data Engineer Secunderabad
🏢 RandomTrees
📍 Secunderabad