Job Description
Role Overview
nWe are looking for a Lead Data Engineer with strong experience building scalable batch, near-real-time, and streaming data platforms on Microsoft Azure.
nThe role requires hands-on expertise in Python, PySpark, advanced SQL, Azure data services, Medallion Architecture, and deploying Apache Spark workloads on Kubernetes or Azure Kubernetes Service (AKS).
n
n.Key Responsibilities
n
- n
- Design and build batch, near-real-time, and streaming data pipelines on Azure.n
- Develop Bronze, Silver, and Gold data layers using Medallion Architecture.n
- Build and deploy containerized PySpark workloads on Kubernetes or AKS.n
- Configure Spark drivers, executors, CPU, memory, scaling, dependencies, and storage access.n
- Integrate data from REST APIs, SFTP, databases, files, enterprise systems, and Azure Event Hubs.n
- Develop complex transformation, cleansing, enrichment, reconciliation, and validation workflows.n
- Implement incremental loads, CDC, watermarking, deduplication, schema evolution, retries, and recovery.n
- Optimize Spark jobs,
partitioning, shuffles, joins, file sizes, and query performance.n
- Implement monitoring, logging, alerting, audit controls, and data-quality checks.n
- Build reusable Python, PySpark, and SQL components.n
- Create CI/CD pipelines for Spark applications, Docker images, and Kubernetes deployments.n
- Review technical designs and support data engineers with implementation standards.n
n
nRequired Skills
n
- n
- 6+ years of hands-on data engineering experience.n
- Robust experience with Microsoft Azure data platforms.n
- Advanced Python, PySpark, and SQL skills.n
- Strong hands-on experience with: Apache Spark, Kubernetes and AKS, Docker, Azure Data Lake Storage Gen2, Azure Event Hubs, Azure DevOps and Gitn
- Experience deploying and operating Spark applications on Kubernetes.n
- Strong understanding of Spark drivers, executors, resource allocation, partitioning, caching, broadcast j
📌 Lead Data Engineer (Pune)
🏢 ORMAE
📍 Pune