- Design, develop, and maintain end-to-end data pipelines using Azure services.
- Build ETL/ELT processes for ingesting, transforming, and loading large-scale data.
- Develop and optimize PySpark applications in Azure Databricks.
- Implement scalable data lake, data warehouse, and lakehouse architectures.
- Integrate data from multiple sources including databases, APIs, files, and cloud platforms.
- Monitor, troubleshoot, and optimize data pipeline performance.
- Ensure data quality, governance, security, and compliance standards.
- Collaborate with Data Scientists, BI teams, Architects, and business stakeholders.
- Implement CI/CD pipelines and DevOps best practices.
- Create technical documentation and support production deployments