20 Aug
|
RandomTrees
|
Secunderabad
20 Aug
RandomTrees
Secunderabad
Role & Responsibilities:
- Design, develop, and maintain ELT/ETL pipelines on GCP using Dataflow/Beam, Dataproc/Spark, and Airflow/Composer.
- Model and optimize datasets in BigQuery using partitioning, clustering, materialized views, and UDFs.
- Build streaming and near-real-time data ingestion using Pub/Sub, Dataflow, and CDC where applicable.
- Implement data quality checks, validation frameworks, and SLAs, monitoring pipelines via Cloud Monitoring/Logging.
- Optimize performance and cost across GCS, Dataproc autoscaling, and BigQuery slot usage.
- Contribute to coding, CI/CD standards, observability, documentation, and perform code reviews.
- Partner with Analytics, BI, and ML teams to productize datasets and ensure strong data contracts.
- Support production operations, including on-call rotations for critical pipelines.
Preferred Candidate Profile:
- Hands-on expertise in GCP data stack: BigQuery, Dataflow (Apache Beam), Dataproc, Cloud Storage, Pub/Sub, Cloud Composer (Airflow).
- Strong experience with Spark (PySpark or Scala) for batch processing.
- Solid Airflow DAG design knowledge (idempotent tasks, backfills, retries, SLAs).
- Advanced SQL and data modeling (star/snowflake schemas, slowly changing dimensions, partition strategies).
- Proficiency in Python (preferred) or Scala/Java.
- Experience with Git and CI/CD tools (Cloud Build, GitHub Actions, GitLab CI).
- Familiarity with GCP security and governance (IAM, service accounts, secrets management, VPC-SC).
- Strong debugging, problem-solving, and communication skills.
Positive-to-Have:
- Snowflake (migration, performance tuning, tasks/streams).
- Terraform for infrastructure as code on GCP.
- Kafka or other streaming tools; Cloud Run/Functions for glue services.
- BI exposure (Looker, Tableau, Power BI).
- GCP Professional Data Engineer certification.
📌 Data Engineer (Secunderabad)
🏢 RandomTrees
📍 Secunderabad