17 Sep
|
Transport Corporation Of India
|
India
17 Sep
Transport Corporation Of India
India
Role: Data Engineer
Job Purpose
Design, build and maintain scalable data pipelines from source systems into Snowflake or Databricks, ensuring performance, reliability and data quality across the pipeline lifecycle.
Key Responsibilities
Design and develop ETL/ELT pipelines using PySpark for batch and incremental loads.
Write complex, performance-tuned SQL for transformations, reconciliation and reporting layers.
Build ingestion pipelines from on-prem/source systems into cloud storage (landing zone), then into Snowflake or Databricks (Raw/Bronze Curated/Silver Presentation/Gold).
Manage Snowflake objects (warehouses, roles/RBAC, Snowpipe, Streams/Tasks, Time Travel) or Databricks (clusters, Delta Lake, Unity Catalog, notebooks).
Implement data quality and reconciliation frameworks across pipeline stages.
Handle schema design, partitioning/clustering, and query/job performance tuning.
Set up and monitor pipeline orchestration (Airflow).
Tech Stack
Languages/Query: Advanced SQL (window functions, query optimization, stored procedures), Python
Processing: PySpark (DataFrame API, partitioning, performance tuning, skew handling)
Data Platform: Snowflake (warehouses, RBAC, Snowpipe, Streams/Tasks) and/or Databricks (clusters, Delta Lake, Unity Catalog)
Cloud: Valuable to have knowledge of any cloud platform (AWS/GCP/Azure)
File Formats: Parquet, Delta, ORC
Orchestration: Airflow/Dagster/ADF/Snowflake Tasks
Other: Git, CI/CD basics, data modeling (dimensional modeling, SCD types)
Qualification: B.Tech/B.E./MCA or equivalent — 2 (0.5) years relevant experience.
📌 Data Engineer Gurugram (India)
🏢 Transport Corporation Of India
📍 India