Lead Data Engineer
Location: Pune
Experience: 6+ Years
Role Overview
We are looking for a Lead Data Engineer with strong experience building scalable batch, near-real-time, and streaming data platforms on Microsoft Azure.
The role requires hands-on expertise in Python, PySpark, advanced SQL, Azure data services, Medallion Architecture, and deploying Apache Spark workloads on Kubernetes or Azure Kubernetes Service (AKS).
Key Responsibilities
Design and build batch, near-real-time, and streaming data pipelines on Azure.
Develop Bronze, Silver, and Gold data layers using Medallion Architecture.
Build and deploy containerized PySpark workloads on Kubernetes or AKS.
Configure Spark drivers, executors, CPU, memory, scaling, dependencies, and storage access.
Integrate data from REST APIs, SFTP, databases, files, enterprise systems, and Azure Event Hubs.
Develop complex transformation, cleansing, enrichment, reconciliation, and validation workflows.
Implement incremental loads, CDC, watermarking, deduplication, schema evolution, retries, and recovery.
Optimize Spark jobs,
partitioning, shuffles, joins, file sizes, and query performance.
Implement monitoring, logging, alerting, audit controls, and data-quality checks.
Build reusable Python, PySpark, and SQL components.
Create CI/CD pipelines for Spark applications, Docker images, and Kubernetes deployments.
Review technical designs and support data engineers with implementation standards.
Required Skills
6+ years of hands-on data engineering experience.
Robust experience with Microsoft Azure data platforms.
Advanced Python, PySpark, and SQL skills.
Strong hands-on experience with:
o Apache Spark
o Kubernetes and AKS
o Docker
o Azure Data Lake Storage Gen2
o Azure Event Hubs
o Azure DevOps and Git
Experience deploying and operating Spark applications on Kubernetes.
Strong understanding of Spark drivers, executors, resource allocation, partitioning, caching, broadcast joins, shuffle optimization,
📌 Data Engineering Lead (Pune)
🏢 ORMAE
📍 Pune