Job Title: Developer
Work Location: Chennai / Bangalore /Hyderabad
Skill Required: Digital : BigData and Hadoop Ecosystems~Digital : Kafka~Digital : Google Cloud
Experience Range: 10+ Years
Big Data with GCP Tech Lead:
- 10+ Years of experience in developing and migrating Big Data applications in to GCP
- Should be very strong in Big Data skill set Pyspark , Hive , Hadoop and HQL
- Hands-on experience with core GCP services relevant to data engineering, including Dataproc, BigQuery, Cloud Storage, Dataflow, Pub/Sub, and IAM.
- Proven expertise in developing, optimizing, and deploying complex PySpark applications for large-scale data processing.
- Create scalable data pipelines and ETL/ELT processes to extract, transform, and load data from various sources into GCP services.
- Build and optimize data models, queries, and storage solutions using GCP tools like BigQuery for efficient data analysis and reporting.
- Utilize tools like Cloud Composer and Apache Airflow to automate data workflows and processes
- Analyze, refactor, and optimize existing PySpark code for effective execution within GCP environments, potentially leveraging services like Dataproc, Dataflow, or BigQuery Spark.
- Design, develop, and deploy scalable and robust data pipelines on GCP using services such as Cloud Storage, BigQuery, Pub/Sub, Dataflow, and Dataproc.
- Oversee and execute the migration of large-scale datasets from source systems to GCP storage solutions (e.g., Cloud Storage, BigQuery).
- Create detailed documentation for migrated systems and establish best practices for PySpark development and data engineering on GCP.
- Excellent analytical and problem-solving skills to diagnose and resolve complex technical issues.
- Should have experience in working in Agile Methodology
- Should be flexible in timings
GCP Engineers
- 7+ Years of experience in developing and migrating Big Data applications in to GCP
- Should be very strong in Big Data skill set Pyspark , Hive , Hadoop and HQL, Kafka and API's
- Hands-on experience with core GCP services relevant to data engineering, including Dataproc, BigQuery, Cloud Storage, Dataflow, Pub/Sub, and IAM, DAG in Cloud Composer
- Proven expertise in developing, optimizing, and deploying complex PySpark applications for large-scale data processing.
- Create scalable data pipelines and ETL/ELT processes to extract, transform, and load data from various sources into GCP services.
- Build and optimize data models, queries, and storage solutions using GCP tools like BigQuery for efficient data analysis and reporting.
- Utilize tools like Cloud Composer and Apache Airflow to automate data workflows and processes
- Analyze, refactor, and optimize existing PySpark code for efficient execution within GCP environments, potentially leveraging services like Dataproc, Dataflow, or BigQuery Spark.
- Design, develop, and deploy scalable and robust data pipelines on GCP using services such as Cloud Storage, BigQuery, Pub/Sub, Dataflow, and Dataproc.
- Oversee and execute the migration of large-scale datasets from source systems to GCP storage solutions (e.g., Cloud Storage, BigQuery).
- Create detailed documentation for migrated systems and establish best practices for PySpark development and data engineering on GCP.
- Excellent analytical and problem-solving skills to diagnose and resolve complex technical issues.
- Should have experience in working in Agile Methodology
- Should be flexible in timings
GCP Engineers –
- 5+ Years of experience in developing and migrating Big Data applications in to GCP
- Hands-on experience with core GCP services relevant to data engineering, including Dataproc, BigQuery, Cloud Storage, Dataflow, Pub/Sub, and IAM, DAG in Cloud Composer
- Create scalable data pipelines and ETL/ELT processes to extract, transform, and load data from various sources into GCP services.
- Build and optimize data models, queries, and storage solutions using GCP tools like BigQuery for efficient data analysis and reporting.
- Utilize tools like Cloud Composer and Apache Airflow to automate data workflows and processes
- Analyze, refactor, and optimize existing PySpark code for efficient execution within GCP environments, potentially leveraging services like Dataproc, Dataflow, or BigQuery Spark.
- Design, develop, and deploy scalable and robust data pipelines on GCP using services such as Cloud Storage, BigQuery, Pub/Sub, Dataflow, and Dataproc.
- Oversee and execute the migration of large-scale datasets from source systems to GCP storage solutions (e.g., Cloud Storage, BigQuery).
- Create detailed documentation for migrated systems and establish best practices for data engineering on GCP.
- Excellent analytical and problem-solving skills to diagnose and resolve complex technical issues.
- Should have experience in working in Agile Methodology
- Should be flexible in timings
📌 GCP Data Engineer (India)
🏢 Clifyx
📍 India