10 Sep
|
Tata Consultancy Services
|
Hyderabad
10 Sep
Tata Consultancy Services
Hyderabad
Job Description
Weekday virtual drive
n
n
7+ years
n
21-Aug-26
n
12-2pm
n
Hyderabad
n
n
We are pleased to invite you for an interview scheduled on
n
Interview Details:• Date:
n
n
n
21-Aug-2
n
12-2pm
n
Hyderabd
n
n
n
Please share updated resume
n
Name:
n
Contact Number:
n
Email ID:
n
Highest Qualification in: (Eg. B.Tech/B.E./M.Tech/MCA/M.Sc./MS/BCA/B.Sc./Etc.)
n
Current Organization Name:
n
Total IT Experience-7 to 10 yrs
n
LOCATION TCS Hyderabab
n
Current CTC
n
Expected CTC
n
Notice period:
n
Whether worked with TCS - Y/N
n
Please apply only if your skill matches
n
n
n
Digital : Python(MongoDB, Python, Pyspark, Big Query, GCS)
n
n
1
n
Role**
n
Mongo db GCP Data Engineer Python Pyspark Developer (BigQuery, Cloud Storage, Dataproc, Airflow)
n
2
n
Required Technical Skill Set**
n
GCP Data Engineer to design, build, and optimize scalable data pipelines and analytics solutions using BigQuery, Cloud Storage, Dataproc, and Airflow.
n
n
n
Desired Experience Range**
n
7+ Years
n
n
Location of Requirement
n
HYDERABAD
n
n
Immediate Joiners Needed
n
n
n
n
n
n
Desired Competencies (Technical/Behavioral Competency)
n
Must-Have**
n
(Ideally should not be more than 3-5)
n
n
- GCP Services: BigQuery, Cloud Storage, Dataproc, Cloud Composer (managed Airflow) or self-managed Airflow.
n
- Airflow: Robust experience in DAG creation, operators/hooks,
scheduling, backfilling, retry strategies, and CI/CD for DAG deployments.
n
- Programming: Proficiency in Python and Pyspark (PySpark, Airflow DAGs), SQL (advanced BigQuery SQL).
n
- Data Modeling: Dimensional modeling (Star/Snowflake), data vault basics, and schema design for analytics.
n
- Performance Tuning: BigQuery partitioning/clustering, predicate pushdown, job stats review, Dataproc executor tuning.
n
- Version Control & CI/CD: Git, branching strategies, pipelines for deploying Airflow DAGs and config.
n
- Operational Excellence: Monitoring with Stackdriver/Cloud Logging, debugging pipeline failures, and root-cause analysis.
n
- involves end-to-end ownership of data ingestion, transformation, orchestration, and performance tuning for batch and near real-time workflows.
n
n
Good-to-Have
n
n
- Streaming: Pub/Sub, Dataflow (Apache Beam) for near real-time pipelines.
n
- Orchestration Patterns: Event-driven pipelines, dependency management, and cross-environment promotion.
n
- Data Governance: Catalog/lineage tools (e.g., Data Catalog), PII handling, row-level security, column-level encryption.
n
- Containers & Infra: Docker, Terraform for IaC on GCP; Kubernetes concepts.
n
- BI Integration: Experience integrating with Looker, Tableau, or Power BI.
n
- Certifications: Google Professional Data Engineer / Cloud Architect.
n
n
📌 Digital: Python(MongoDB, Python, Pyspark, Big Query, GCS) (Hyderabad)
🏢 Tata Consultancy Services
📍 Hyderabad