- Design, develop, and optimize scalable data pipelines using AWS services.
- Build and maintain batch and real-time data ingestion and processing frameworks.
- Develop enterprise-grade data warehousing solutions using Py-Spark.
- Convert Scala-based ETL, batch, and streaming pipelines into PySpark frameworks.
- Optimize PySpark jobs for performance, scalability, and resource utilization.
- Support cloud modernization initiatives on AWS/Databricks platforms.