Key Responsibilities
- Lead the
modernization and migration of legacy data applications, ETL pipelines, and data platforms to Databricks
.
- Assess existing applications, data pipelines, databases, and workloads and define appropriate migration strategies.
- Design scalable
Databricks Lakehouse architectures
using Delta Lake, Unity Catalog, Databricks SQL, and Apache Spark.
- Develop and optimize
PySpark / Spark SQL
workloads for large-scale data processing.
- Modernize legacy ETL/data-processing workloads and convert them into Databricks-native pipelines.
- Design
Bronze, Silver, and Gold / Medallion architecture
for enterprise data platforms.
- Work on batch and near-real-time data ingestion and transformation pipelines.
- Implement data governance, security, access control, lineage, and data quality using
Unity Catalog
.
- Optimize Databricks jobs for
performance, scalability, reliability, and cost
.
- Support migration from traditional data warehouses, Hadoop, on-premises platforms, or legacy ETL tools to Databricks.
- Define migration frameworks, reusable patterns,
coding standards, and best practices.
- Collaborate with application, cloud, DevOps, data engineering, and business teams.
- Perform technical POCs and evaluate modernization approaches.
- Troubleshoot complex Spark/Databricks performance and production issues.
- Provide technical mentoring and guidance to data engineering teams.
Mandatory Technical Skills
- 10+ years of experience
in Data Engineering / Data Architecture.
- Solid hands-on experience with
Databricks
.
- Excellent knowledge of
Apache Spark, PySpark, and Spark SQL
.
- Strong experience with
Delta Lake
and Lakehouse architecture.
- Hands-on experience with
Unity Catalog
.
- Strong SQL and database concepts.
- Experience in
ETL/ELT pipeline development and modernization
.
- Experience migrating legacy/on-premises data workloads to cloud platforms.
- Strong understanding of data modeling and data architecture.
- Experience with
Databrick
📌 Lead DataBricks Engineer (Pune)
🏢 Atyeti
📍 Pune