16 Aug
|
Tata Consultancy Services
|
Hyderabad
16 Aug
Tata Consultancy Services
Hyderabad
Information Details 1 Role GCP Data Engineer Postgres Python Pyspark Developer (Big Query, Cloud Storage, Dataproc, Airflow) 2 Required Technical Skill Set GCP Data Engineer to design, build, and optimize scalable data pipelines and analytics solutions using Big Query, Cloud Storage, Dataproc, and Airflow. 3 No of Requirements 5 4 Desired Experience Range 7+ Years 5 Location of Requirement HYDERABAD 6 Immediate Joiners Needed YES Desired Competencies (Technical/Behavioral Competency) Must-Have** (Ideally should not be more than 3-5) · GCP Services: Big Query, Cloud Storage, Dataproc, Cloud Composer (managed Airflow) or self-managed Airflow. · Airflow: Robust experience in DAG creation, operators/hooks, scheduling, backfilling, retry strategies, and CI/CD for DAG deployments. · Programming: Proficiency in Postgres Programming including Tables, Triggers, Views, Stored procedures, Python and Pyspark (PySpark, Airflow DAGs), SQL (advanced BigQuery SQL). · Data Modeling: Dimensional modeling (Star/Snowflake), data vault basics, and schema design for analytics. · Performance Tuning: BigQuery partitioning/clustering, predicate pushdown, job stats review, Dataproc executor tuning. · Version Control &
• CI/CD: Git, branching strategies, pipelines for deploying Airflow DAGs and config. · Operational Excellence: Monitoring with Stackdriver/Cloud Logging, debugging pipeline failures, and root-cause analysis. · involves end-to-end ownership of data ingestion, transformation, orchestration, and performance tuning for batch and near real-time workflows. Good-to-Have · Streaming: Pub/Sub, Dataflow (Apache Beam) for near real-time pipelines. · Orchestration Patterns: Event-driven pipelines, dependency management,
and cross-environment promotion. · Data Governance: Catalog/lineage tools (e.g., Data Catalog), PII handling, row-level security, column-level encryption. · Containers &
• Infra: Docker, Terraform for IaC on GCP
• Kubernetes concepts. · BI Integration: Experience integrating with Looker, Tableau, or Power BI. · Certifications: Google Professional Data Engineer / Cloud Architect. Responsibility of / Expectations from the Role 1 Data Pipeline Development: Build robust ETL/ELT pipelines using Apache Airflow (DAG creation, scheduling, monitoring) to orchestrate data workflows across GCP services. 2 Data Warehousing: Design and optimize BigQuery schemas, partitioning/clustering strategies, materialized views, and query performance tuning. 3 Data Processing: Implement scalable data processing using Dataproc (Spark/Hive), including job configuration, optimization, and cost control. 4 Data Ingestion &
• Storage: Manage ingestion from diverse sources (APIs, files, streaming) and design storage strategies using Cloud Storage (lifecycle policies, tiers, security). 5 Quality &
• Observability: Implement data validation (e.g., Great Expectations or custom checks), logging/alerting, and SLA monitoring for pipelines. 6 Security &
• Governance: Apply IAM, service accounts, VPC SC, CMEK, and access policies across GCP resources; ensure compliance with data governance standards. 7 Cost &
• Performance: Optimize queries, cluster usage, and storage to balance cost/performance; leverage reservation/flex slots and job-level optimizations. 8 Collaboration: Work closely with analytics, product, and business teams to translate requirements into scalable data solutions; create documentation and handover materials.
📌 GCP Data Engineer Postgres Python Pyspark Developer (Hyderabad)
🏢 Tata Consultancy Services
📍 Hyderabad