12 Aug
|
Tata Consultancy Services
|
Hyderabad
12 Aug
Tata Consultancy Services
Hyderabad
Information Details
1 Role GCP Data Engineer Postgres Python Pyspark Developer (Big Query, Cloud Storage, Dataproc, Airflow)
2 Required Technical Skill Set GCP Data Engineer to design, build, and optimize scalable data pipelines and analytics solutions using Big Query, Cloud Storage, Dataproc, and Airflow.
3 No of Requirements 5
4 Desired Experience Range 7+ Years
5 Location of Requirement HYDERABAD
6 Immediate Joiners Needed YES Desired Competencies (Technical/Behavioral Competency)
Must-Have** (Ideally should not be more than 3-5) · GCP Services: Big Query, Cloud Storage, Dataproc, Cloud Composer (managed Airflow) or self-managed Airflow. · Airflow: Strong experience in DAG creation, operators/hooks, scheduling, backfilling, retry strategies, and CI/CD for DAG deployments. · Programming: Proficiency in Postgres Programming including Tables, Triggers, Views, Stored procedures, Python and Pyspark (PySpark, Airflow DAGs), SQL (advanced BigQuery SQL). · Data Modeling: Dimensional modeling (Star/Snowflake), data vault basics, and schema design for analytics. ·
Performance Tuning: BigQuery partitioning/clustering, predicate pushdown, job stats review, Dataproc executor tuning. · Version Control & CI/CD: Git, branching strategies, pipelines for deploying Airflow DAGs and config. · Operational Excellence: Monitoring with Stackdriver/Cloud Logging, debugging pipeline failures, and root-cause analysis. · involves end-to-end ownership of data ingestion, transformation, orchestration, and performance tuning for batch and near real-time workflows. Good-to-Have · Streaming: Pub/Sub, Dataflow (Apache Beam) for near real-time pipelines.
· Orchestration Patterns: Event-driven pipelines, dependency management,
and cross-setting promotion. · Data Governance: Catalog/lineage tools (e.g., Data Catalog), PII handling, row-level security, column-level encryption. · Containers & Infra: Docker, Terraform for IaC on GCP; Kubernetes concepts.
· BI Integration: Experience integrating with Looker, Tableau, or Power BI. · Certifications: Google Professional Data Engineer / Cloud Architect. Responsibility of / Expectations from the Role
1 Data Pipeline Development: Build robust ETL/ELT pipelines using Apache Airflow (DAG creation, scheduling, monitoring) to orchestrate data workflows across GCP services.
2 Data Warehousing: Design and optimize BigQuery schemas, partitioning/clustering strategies, materialized views, and query performance tuning.
3 Data Processing: Implement scalable data processing using Dataproc (Spark/Hive), including job configuration, optimization, and cost control.
4 Data Ingestion & Storage: Manage ingestion from diverse sources (APIs, files, streaming) and design storage strategies using Cloud Storage (lifecycle policies, tiers, security).
5 Quality & Observability: Implement data validation (e.g., Great Expectations or custom checks), logging/alerting, and SLA monitoring for pipelines.
6 Security & Governance: Apply IAM, service accounts, VPC SC, CMEK, and access policies across GCP resources; ensure compliance with data governance standards.
7 Cost & Performance: Optimize queries, cluster usage, and storage to balance cost/performance; leverage reservation/flex slots and job-level optimizations.
8 Collaboration: Work closely with analytics, product, and business teams to translate requirements into scalable data solutions; create documentation and handover materials.
📌 GCP Data Engineer Postgres Python Pyspark Developer (Hyderabad)
🏢 Tata Consultancy Services
📍 Hyderabad