• Azure DataBricks/Spark (Mandatory )
- Azure Data Factory (Mandatory)
- Azure Analysis Services (Useful)
- Azure ML (Nice to have) • Azure Cosmos DB (Helpful)
- Azure EventHubs (Useful) • PowerBI (Useful)
Good-to-Have
- Experience in Databricks or other cloud-based data platforms.
- Exposure to Terraform or Infrastructure-as-Code (IaC) for data platform automation.
• Certification in Azure Data Engineering (DP-203) or relevant credentials.
- Familiarity with Machine Learning and AI-driven data processing.
• Knowledge of data lake architecture and best practices.
- Exposure to containerization technologies like Docker and Kubernetes.
- Knowledge of Python, Spark, or other scripting languages is a plus
Responsibility of / Expectations from the Role
- Responsible for both developing and supporting data products alongside existing data teams
- Deliver outcomes and as well as support products in production (80/20 split)
- Design, develop, and maintain scalable ETL/ELT pipelines using Azure Data Factory (ADF), Azure Synapse, and DBT.
• Implement data transformation, data modeling, and data quality best practices.
- Develop and optimize SQL-based transformations in Azure Synapse.
• Ensure data integrity, security, and governance in compliance with best practices.
- Collaborate with data analysts, data scientists, and business stakeholders to understand
data requirements.
- Work towards automating and streamlining data engineering workflows.
- Optimize cloud data infrastructure for performance and cost efficiency.
• Support and maintain CI/CD pipelines for data engineering processes.