Role OverviewWe are looking for a Lead Data Engineer with strong experience building scalable batch, near-real-time, and streaming data platforms on Microsoft Azure.The role requires hands-on expertise in Python, PySpark, advanced SQL, Azure data services, Medallion Architecture, and deploying Apache Spark workloads on Kubernetes or Azure Kubernetes Service (AKS).
.Key ResponsibilitiesDesign and build batch, near-real-time, and streaming data pipelines on Azure.Develop Bronze, Silver, and Gold data layers using Medallion Architecture.Build and deploy containerized PySpark workloads on Kubernetes or AKS.Configure Spark drivers, executors, CPU, memory, scaling, dependencies, and storage access.Integrate data from REST APIs, SFTP, databases, files, enterprise systems, and Azure Event Hubs.Develop complex transformation, cleansing, enrichment, reconciliation, and validation workflows.Implement incremental loads, CDC, watermarking, deduplication, schema evolution, retries, and recovery.Optimize Spark jobs, partitioning, shuffles, joins, file sizes, and query performance.Implement monitoring, logging, alerting, audit controls, and data-quality checks.Build reusable Python, PySpark, and SQL components.Create CI/CD pipelines for Spark applications, Docker images,
and Kubernetes deployments.Review technical designs and support data engineers with implementation standards.
Required Skills6+ years of hands-on data engineering experience.Strong experience with Microsoft Azure data platforms.Advanced Python, PySpark, and SQL skills.Robust hands-on experience with: Apache Spark, Kubernetes and AKS, Docker, Azure Data Lake Storage Gen2, Azure Event Hubs, Azure DevOps and GitExperience deploying and operating Spark applications on Kubernetes.Strong understanding of Spark drivers, executors, resource allocation, partitioning, caching, broadcast joins, shuffle optimization, and skew handling.Experience with batch, streaming, ETL, ELT, and event-driven processing patterns.Experience implementing Medallion Architecture.Experience with REST APIs, SFTP, JSON, CSV, Parquet, Delta Lake, and relational databases.Experience with CDC, incremental processing, schema enforcement, and schema evolution.Strong understanding of data modelling, schema design, partitioning, and storage optimization.Experience with Kubernetes Jobs, Cron Jobs, Config Maps, Secrets, resource limits, node pools, and autoscaling.Experience implementing pipeline observability, data validation, monitoring, alerting, and error recovery.
📌 Lead Data Engineer (Mumbai)
🏢 ORMAE
📍 Mumbai