Job Overview
We are looking for a Lead Data Engineer with solid expertise in Databricks, PySpark, Python, and SQL to design, develop, and maintain scalable data pipelines. The role involves modernizing legacy ETL processes, implementing Delta Lake solutions, optimizing Spark workloads, mentoring engineers, and collaborating with cross-functional teams to deliver enterprise-scale data engineering solutions.
Key Responsibilities
Data Pipeline Development & Operations
Design, build, and operate scalable and reliable data pipelines on the Databricks platform.
Develop end-to-end data workflows from ingestion through transformation to consumption.
Implement robust error handling, monitoring, and alerting mechanisms.
Ensure data pipeline reliability, performance, and maintainability.
Optimize Spark job performance and cluster configurations.
Manage and orchestrate Databricks Jobs and workflows.
Legacy Code Modernization
Refactor legacy code and data pipelines to PySpark.
Migrate traditional ETL processes to contemporary ELT patterns on Databricks.
Assess existing codebases for modernization opportunities.
Ensure data integrity and backward compatibility during migrations.
Document migration approaches and refactoring playbooks.
Collaborate with stakeholders during transition activities.
Data Engineering Excellence
Implement data quality checks and validation frameworks.
Design and maintain Delta Lake tables with optimization strategies.
Develop reusable code libraries and frameworks.
Follow software engineering best practices including Git, testing, and CI/CD.
Participate in code reviews.
Troubleshoot production data pipeline issues.
Collaboration & Leadership
Work closely with data architects, analysts, and business stakeholders.
Collaborate with Infrastructure, Applications, and Cyber teams.
Share technical knowledge and best practices.
Mentor junior data engineers on Databricks and PySpark.
Maintain technical documentation.
📌 Lead Data Engineer Databricks & Pyspark Pune
🏢 Experis
📍 Pune