19 Sep
|
Orion Systems
|
Noida
19 Sep
Orion Systems
Noida
Role & responsibilities
- Design, build, and deploy robust ETL/ELT pipelines within the lakehouse platform (Google BigQuery or Databricks) using SQL, Python, PySpark, and Spark SQL.
- Implement and manage the Medallion Architecture (Bronze, Silver, Gold layers) using Delta Lake or BigQuery datasets to ensure data quality and progressive data refinement.
- Leverage native ingestion tooling such as BigQuery Data Transfer Service, Pub/Sub streaming, or Databricks Auto Loader for efficient, scalable, and incremental ingestion of data from sources like GA4 into the Bronze layer.
- Develop, schedule, and monitor complex, multi-task data workflows using Cloud Composer (Airflow), BigQuery scheduled queries, or Databricks Workflows.
- Optimize BigQuery tables (partitioning, clustering, materialised views) and Spark jobs / Delta Lake tables (using techniques like OPTIMIZE, Z-ORDER, and partitioning) for high performance and cost efficiency.
- Implement data governance, security, and discovery using Dataplex / BigQuery policy tags or Unity Catalog, including managing access controls and data lineage.
- Write complex, customized SQL queries to manipulate data and support ad-hoc analytical requests from business teams.
- Develop strategies for data ingestion from multiple sources, using various techniques including streaming, API consumption, and replication.
- Document data engineering processes, data models, and technical specifications for
the lakehouse platform.
- Conform to agile development practices, including version control (Git), continuous
integration/delivery (CI/CD), and test-driven development.
- Provide production support for data pipelines, actively monitoring and resolving
issues to ensure the continuous flow of critical data.
- Collaborate with analytics and business teams to understand data requirements and
deliver well-modelled, performant datasets in the gold layer.
Preferred candidate profile
- Lakehouse Platform Expertise (Google BigQuery and/or Databricks):
- BigQuery: Deep, hands-on experience with BigQuery architecture, including
partitioning, clustering, materialised views, slot/cost optimisation, and diagnosing query performance using query plans and INFORMATION_SCHEMA.
- Apache Spark / Delta Lake: Strong experience with Spark architecture, writing and
optimising PySpark and Spark SQL jobs, and building reliable pipelines on Delta Lake. Proficient with ACID transactions, time travel, schema evolution, and DML operations (MERGE, UPDATE, DELETE).
- Data Ingestion: Experience with modern ingestion tools, such as BigQuery Data
Transfer Service, Pub/Sub / Dataflow streaming, Databricks Auto Loader, and COPY INTO for scalable file processing.
- Data Governance: Robust understanding of data governance concepts and practical
experience implementing security, lineage, and discovery using Dataplex, BigQuery IAM and policy tags, or Unity Catalog.
📌 Senior Data Engineer : GCP (Noida)
🏢 Orion Systems
📍 Noida