- Design, develop, and optimize scalable data pipelines using scala and Pyspark in AWS setting.
- Build and maintain batch and real-time data ingestion and processing frameworks.
- Develop enterprise-grade data warehousing solutions using Scala and Pyspark.
- Analyze existing Scala and Spark applications and identify migration requirements.
- Convert Scala-based ETL, batch, and streaming pipelines into PySpark frameworks.
- Optimize PySpark jobs for performance, scalability, and resource utilization.
- Support cloud modernization initiatives on AWS/Databricks/Snowflake platforms.
- Implement ETL/ELT processes for structured and unstructured data.