22 Sep
|
Transport Corporation Of India
|
Gurugram
22 Sep
Transport Corporation Of India
Gurugram
Role: Data Engineer
Job Purpose
Design, build and maintain scalable data pipelines from source systems into Snowflake or Databricks, ensuring performance, reliability and data quality across the pipeline lifecycle.
Key Responsibilities
- Design and develop ETL/ELT pipelines using PySpark for batch and incremental loads.
- Write complex, performance-tuned SQL for transformations, reconciliation and reporting layers.
- Build ingestion pipelines from on-prem/source systems into cloud storage (landing zone), then into Snowflake or Databricks (Raw/Bronze Curated/Silver Presentation/Gold).
- Manage Snowflake objects (warehouses, roles/RBAC, Snowpipe, Streams/Tasks, Time Travel) or Databricks (clusters, Delta Lake, Unity Catalog, notebooks).
- Implement data quality and reconciliation frameworks across pipeline stages.
- Handle schema design, partitioning/clustering, and query/job performance tuning.
- Set up and monitor pipeline orchestration (Airflow).
Tech Stack
- Languages/Query: Advanced SQL (window functions, query optimization, stored procedures), Python
- Processing: PySpark (DataFrame API, partitioning, performance tuning, skew handling)
- Data Platform: Snowflake (warehouses, RBAC, Snowpipe, Streams/Tasks) and/or Databricks (clusters, Delta Lake, Unity Catalog)
- Cloud: Valuable to have knowledge of any cloud platform (AWS/GCP/Azure)
- File Formats: Parquet, Delta, ORC
- Orchestration: Airflow/Dagster/ADF/Snowflake Tasks
- Other: Git, CI/CD basics, data modeling (dimensional modeling, SCD types)
Qualification: B.Tech/B.E./MCA or equivalent — 2 (0.5) years relevant experience.
📌 Data Engineer (Gurugram)
🏢 Transport Corporation Of India
📍 Gurugram