16 Aug
|
Important Business
|
Mumbai
16 Aug
Important Business
Mumbai
About the Role
We are seeking a highly skilled Data Engineer with solid expertise in Azure Databricks, SQL, PySpark, and Data Modeling. The ideal candidate will have experience designing and implementing scalable data pipelines, optimizing data workflows, and building modern data platforms on the Azure ecosystem.
*Key Responsibilities
- Design, develop, and maintain ETL/ELT pipelines using Azure Databricks & PySpark.
- Build and manage Delta Lakehouse solutions including Bronze, Silver, Gold layers.
- Collaborate with data architects, analysts, and business stakeholders to design data models (Star/Snowflake schemas, Fact & Dimension tables).
- Optimize Databricks clusters, jobs, and queries for performance and cost-efficiency.
- Implement CI/CD pipelines for Databricks notebooks and data workflows.
- Manage schema evolution, data governance, and quality checks across pipelines.
- Work with Azure Data Lake Storage (ADLS), Azure Synapse Analytics, and SQL Databases for end-to-end data solutions.
- Implement data partitioning, caching, and broadcast joins to optimize PySpark jobs.
- Ensure best practices in data security, compliance, and access management.
- Troubleshoot and optimize slow SQL queries, indexes, and data warehouse performance.
- Support business reporting and analytics needs by designing and maintaining scalable data models.
Requirements
*Required Skills
- Azure Databricks: Notebooks, Delta Tables, Auto-scaling, Job Orchestration.
- SQL: Joins, Window Functions, Indexing, Query Optimization, CTEs, SCD handling.
- PySpark: RDD, Data Frame API, Lazy evaluation, Transformations, Optimizations.
- Data Modeling: OLTP vs OLAP, Star & Snowflake Schema, Fact/Dimension Tables, Normalization/Denormalization.
- Azure Ecosystem: ADLS, Synapse Analytics, Azure Data Factory (ADF) is a plus.
- Strong understanding of ETL best practices, data quality frameworks, and large-scale distributed data processing.
*Nice-to-Have Skills*
- Experience with DataBricks CI/CD pipelines (Azure DevOps/GitHub).
- Familiarity with Power BI/Tableau reporting dashboards.
- Knowledge of Kafka, Event Hub, or real-time streaming.
- Exposure to machine learning pipelines within Databricks.
*Qualifications
- Bachelor’s or Master’s degree in Computer Science, Information Technology, or related field.
- 6-7 years of experience in Data Engineering.
- Proven track record of building scalable data solutions using Azure Databricks and PySpark.
📌 Data Engineer Azure Databricks (Mumbai)
🏢 Important Business
📍 Mumbai