- Must have: Design and develop scalable data pipelines using AWS and Databricks.
- Build and maintain ETL/ELT frameworks using Python and PySpark.
- Process and transform large volumes of structured and unstructured data.
- Develop data ingestion solutions from multiple source systems.
- Optimize Spark jobs and Databricks workloads for performance and cost efficiency.
- Design and implement data lake solutions on AWS.
- Develop reusable data processing frameworks and libraries.
- Collaborate with business and analytics teams to understand data requirements.
- Implement data quality, governance, and security standards.
- Support production deployments and troubleshoot data-related issues.