Lead Data Engineer – PySpark (India)

Lead Data Engineer – PySpark (India)

06 Sep
|
Sparix Global
|
India

06 Sep

Sparix Global

India

Experience:- 8-10 Years

Work Type- Hybrid

Duration:6-12 Months

Shift Timing: 1:00 PM – 10:00 PM

Job Locations:--Bengaluru, India / Pune, India / Noida, India / Gurgaon, India

Skills Required

Primary Skills

PySpark and Databricks

Job Description

Shift Timing: 1:00 PM – 10:00 PM

Key Responsibilities

A. Data Engineering & Pipeline Development

• Design, develop, and optimize large-scale ETL/ELT pipelines using PySpark on Databricks, processing structured and unstructured data at scale.

• Build and maintain Lakehouse architecture (Bronze/Silver/Gold medallion layers) using Delta Lake, ensuring reliability, scalability, and schema evolution support.

• Develop reusable, parameterized, metadata-driven pipeline frameworks for ingestion, transformation, and curation of data from diverse source systems.

• Optimize Spark jobs for performance and cost (partitioning, caching, cluster sizing/auto-scaling, Photon engine, Z-ordering, file compaction).

• Implement data quality checks, validation rules, and monitoring/alerting to ensure pipeline reliability and data trust.





B. Databricks Platform & Cloud Engineering

• Configure and manage Databricks workspaces, clusters, jobs, and workflows; tune cluster policies for cost and performance.

• Implement data governance, access control, and lineage using Unity Catalog.

• Integrate Databricks with cloud services such as Azure Data Lake Storage (ADLS Gen2), Azure Data Factory, Event Hub/Kafka, and Key Vault (or AWS S3/Glue/EMR equivalents).

• Build and maintain CI/CD pipelines for Databricks notebooks, jobs, and Delta Live Tables using Azure DevOps/GitHub Actions and Databricks Repos.

C. Data Modeling, Architecture & Governance

• Design and maintain data models (dimensional, medallion, canonical) to support analytics, reporting, and downstream consumption.

• Partner with Data Architects on target-state data platform design, migration strategy, and platform standards.

• Ensure adherence to data governance, security, a

📌 Lead Data Engineer – PySpark (India)
🏢 Sparix Global
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: lead data engineer – pyspark (india) / india

Subscribe to this job alert:

Get the latest job offers by email for: lead data engineer – pyspark (india) / india