Data Pipeline Development
- Design, develop, and optimize scalable batch and streaming data pipelines using Databricks.
- Build robust ETL/ELT frameworks using PySpark and Spark SQL.
- Develop ingestion pipelines from:
- APIs
- Kafka
- Cloud Storage
- Databases
- Event-driven sources
Lakehouse Architecture
- Design and maintain Medallion Architecture:
- Bronze Layer
- Silver Layer
- Gold Layer
- Implement Delta Lake features including:
- ACID Transactions
- Schema Evolution
- MERGE
- UPSERT
- Time Travel
Data Engineering & Modeling
- Build dimensional and Lakehouse data models.
- Support analytical and reporting workloads.
- Ensure data quality, reconciliation, lineage, and governance.