Role & responsibilities
- Design, develop, and implement scalable Enterprise Data Warehouse (EDW) and Data Lakehouse solutions on Microsoft Azure using Azure Databricks and Delta Lake.
- Build, optimize, and maintain enterprise-grade ETL/ELT pipelines to ingest, transform, and process data from multiple sources, including relational databases, APIs, flat files, and other structured/unstructured data sources.
- Develop high-performance data transformation solutions using Python, PySpark, and Spark SQL within Azure Databricks.
- Lead the modernization and migration of legacy data warehouse environments, SQL stored procedures, Unix scripts, and traditional ETL workflows to cloud-native Azure platforms.
- Develop and orchestrate data pipelines using Azure services such as Azure Data Factory (ADF), Azure Data Lake Storage Gen2 (ADLS Gen2), Azure SQL/Synapse, and Azure Key Vault.
- Monitor, troubleshoot, and optimize Spark jobs, SQL queries, and ETL processes to improve performance, scalability, and cost efficiency.
- Design and implement robust data models and warehouse architectures following industry best practices.
- Collaborate with architects, business stakeholders, data scientists, and cross-functional teams to deliver scalable, secure, and reliable data solutions.
- Lead technical discussions, mentor data engineering teams, conduct code reviews, and promote engineering best practices including CI/CD, automated testing, and coding standards.
- Ensure data quality, governance, security, and compliance standards are incorporated into all data engineering solutions.
Preferred candidate profile
- 10-12 years of experience in Data Engineering, Business Intelligence, or Enterprise Data Warehousing.
- Robust experience designing and implementing enterprise-scale Data Warehouse and Data Lakehouse solutions.
- Hands-on expertise in Azure Databricks with at least 4 years of experience developing production-grade applications using PySpark, Spark SQL, and Delta Lake.
- Stron
📌 Data Engineer (Bengaluru)
🏢 CGI
📍 Bengaluru