Role & responsibilities
- Design, develop, and implement contemporary, scalable Data Warehouse and Data Lakehouse solutions on Microsoft Azure using Databricks and Delta Lake.
- Lead the architecture and development of enterprise-grade ETL/ELT pipelines for ingesting, transforming, and loading large-scale data from multiple sources, including databases, APIs, and flat files.
- Develop and optimize production-ready Python and PySpark code in Azure Databricks to support complex data transformations and analytics workloads.
- Drive modernization initiatives by migrating legacy Enterprise Data Warehouse (EDW) environments, stored procedures, Unix scripts, and traditional ETL processes to Azure-based cloud platforms.
- Leverage Azure services such as Azure Data Factory (ADF), Azure Data Lake Storage Gen2 (ADLS Gen2), Azure Key Vault, Azure SQL, and Azure Synapse to build robust data solutions.
- Monitor, troubleshoot, and optimize Spark workloads, SQL queries, and data pipelines to improve performance, scalability, reliability, and cost efficiency.
- Collaborate with enterprise architects, business stakeholders, and data science teams to align data solutions with business and analytical requirements.
- Implement and promote engineering best practices including CI/CD, automated testing, code reviews, and DevOps methodologies.
- Lead and mentor a team of data engineers, providing technical guidance and resolving complex engineering challenges.
- Ensure data governance, security, compliance, and data quality standards are incorporated into all data engineering solutions.
Preferred candidate profile
- 10 to 12 years of experience in Data Engineering, Business Intelligence, or Enterprise Data Warehousing.
- Strong expertise in Data Warehouse design and data modeling, including Kimball, Inmon, Data Vault methodologies, Star Schema, Snowflake Schema, and Slowly Changing Dimensions (SCD Types 1, 2, and 3).
- Proven experience designing and implementing large-scale ETL/ELT frameworks and complex data integration solutions.
- Advanced programming skills in Python, including core Python, Object-Oriented Programming (OOP), Pandas, automation, and API integration.
- Minimum 4+ years of hands-on experience with Azure Databricks, Apache Spark (PySpark and Spark SQL), and Delta Lake.
- Strong knowledge of the Microsoft Azure Data Platform, specifically Azure Data Factory (ADF) and Azure Data Lake Storage Gen2 (ADLS Gen2).
- Expert-level SQL skills with experience in writing, optimizing, and troubleshooting complex queries, window functions, and database performance tuning.
- Demonstrated experience leading cloud migration projects from on-premises data platforms to modern cloud-based architectures.
- Strong leadership, stakeholder management, and team mentoring capabilities.
- Excellent problem-solving, analytical thinking, and communication skills.
Preferred Qualifications
- Experience migrating legacy environments such as SQL Server, Oracle, Teradata, and Unix Shell scripting workloads to Azure cloud platforms.
- Hands-on experience with Azure DevOps, Git, GitHub Actions, and CI/CD implementation.
- Knowledge of Data Governance, Data Quality frameworks, Role-Based Access Control (RBAC), and Unity Catalog in Databricks.
- Experience working within enterprise-scale data platforms and modern Lakehouse architectures.
- Relevant certifications such as
- Microsoft Azure Data Engineer Associate (DP-203)
- Databricks Certified Data Engineer Professional
📌 Azure Databricks Lead (Bengaluru)
🏢 CGI
📍 Bengaluru