Key Responsibilities
Lead the
modernization and migration of legacy data applications, ETL pipelines, and data platforms to Databricks
.
Assess existing applications, data pipelines, databases, and workloads and define appropriate migration strategies.
Design scalable
Databricks Lakehouse architectures
using Delta Lake, Unity Catalog, Databricks SQL, and Apache Spark.
Develop and optimize
PySpark / Spark SQL
workloads for large-scale data processing.
Modernize legacy ETL/data-processing workloads and convert them into Databricks-native pipelines.
Design
Bronze, Silver, and Gold / Medallion architecture
for enterprise data platforms.
Work on batch and near-real-time data ingestion and transformation pipelines.
Implement data governance, security, access control, lineage, and data quality using
Unity Catalog
.
Optimize Databricks jobs for
performance, scalability, reliability, and cost
.
Support migration from traditional data warehouses, Hadoop, on-premises platforms, or legacy ETL tools to Databricks.
Define migration frameworks, reusable patterns,
coding standards, and best practices.
Collaborate with application, cloud, DevOps, data engineering, and business teams.
Perform technical POCs and evaluate modernization approaches.
Troubleshoot complex Spark/Databricks performance and production issues.
Provide technical mentoring and guidance to data engineering teams.
Mandatory Technical Skills
10+ years of experience
in Data Engineering / Data Architecture.
Robust hands-on experience with
Databricks
.
Excellent knowledge of
Apache Spark, PySpark, and Spark SQL
.
Solid experience with
Delta Lake
and Lakehouse architecture.
Hands-on experience with
Unity Catalog
.
Robust SQL and database concepts.
Experience in
ETL/ELT pipeline development and modernization
.
Experience migrating legacy/on-premises data workloads to cloud platforms.
Strong understanding of data modeling and data architecture.
Experience with
Databrick
📌 Lead Databricks Engineer Pune
🏢 Atyeti
📍 Pune