21 Aug
|
RandomTrees
|
Secunderabad
21 Aug
RandomTrees
Secunderabad
Roles & Responsibilities:
Design, develop, and maintain ELT/ETL pipelines on GCP using Dataflow (Apache Beam), Dataproc/Spark, and Cloud Composer (Airflow).
Model and optimize datasets in BigQuery, including partitioning, clustering, materialized views, and UDFs.
Build streaming and near-real-time ingestion using Pub/Sub, Dataflow, and CDC where applicable.
Implement data quality checks, validation frameworks, and SLAs; monitor pipelines via Cloud Monitoring/Logging.
Optimize performance and cost across GCS, Dataproc autoscaling, and BigQuery slot usage.
Contribute to coding standards, CI/CD processes, observability, and documentation; perform code reviews.
Partner with Analytics, BI, and ML teams to productize datasets and ensure strong data contracts.
Support production operations, including on-call rotations for critical pipelines as needed.
Preferred Candidate Profile:
Hands-on expertise in GCP data stack: BigQuery, Dataflow (Apache Beam), Dataproc, Cloud Storage, Pub/Sub, Cloud Composer (Airflow).
Solid Spark skills (PySpark or Scala) for batch processing on Dataproc.
Solid Airflow DAG design including idempotent tasks, retries, backfills, and SLAs.
Advanced SQL and data modeling: star/snowflake schemas, slowly changing dimensions, partition strategies.
Proficiency in Python, Scala, or Java for data engineering.
Experience with Git and CI/CD (Cloud Build, GitHub Actions, GitLab CI).
Familiarity with GCP security and governance (IAM, service accounts, secrets management, VPC-SC basics).
Strong debugging skills, ownership mindset, and ability to communicate clearly with technical and non-technical stakeholders.
Valuable-to-Have:
Experience with Snowflake, dbt, Excellent Expectations, Terraform on GCP.
Knowledge of Kafka or other streaming tools, Cloud Run/Functions.
BI tool exposure: Looker, Tableau, Power BI.
GCP Professional Data Engineer certification.
📌 Data Engineer Secunderabad
🏢 RandomTrees
📍 Secunderabad