We are looking for a Senior Data Engineer with robust experience in Databricks, PySpark, and Python to design, develop, and maintain scalable data pipelines. The role involves modernizing legacy ETL processes, implementing Delta Lake solutions, optimizing Spark workloads, and collaborating with cross-functional teams to deliver reliable data engineering solutions.
Key Responsibilities
Data Pipeline Development & Operations
Design, build, and operate scalable data pipelines on the Databricks platform.
Develop end-to-end data workflows from ingestion to transformation and consumption.
Implement monitoring, alerting, and error-handling mechanisms.
Optimize Spark jobs and cluster configurations.
Manage Databricks Jobs and workflow orchestration.
Legacy Code Modernization
Refactor legacy data pipelines into PySpark.
Migrate ETL processes to ELT on Databricks.
Assess existing codebases for optimization prospects.
Ensure data integrity and backward compatibility.
Document migration approaches and playbooks.
Data Engineering
Implement data quality checks and validation frameworks.
Design and maintain Delta Lake tables.
Develop reusable code libraries.
Follow version control, testing, and CI/CD practices.
Participate in code reviews.
Troubleshoot production data pipelines.
Collaboration
Work with data architects, analysts, business stakeholders, Infrastructure, Applications, and Cyber teams.
Mentor junior engineers on Databricks and PySpark.
Share technical knowledge and maintain documentation.
📌 Senior Data Engineer Databricks & Pyspark Pune
🏢 Experis
📍 Pune
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.