Role & responsibilities
Design, develop, and optimize scalable data pipelines using scala and Pyspark in AWS setting.
Build and maintain batch and real-time data ingestion and processing frameworks.
Develop enterprise-grade data warehousing solutions using Scala and Pyspark.
Analyze existing Scala and Spark applications and identify migration requirements.
Convert Scala-based ETL, batch, and streaming pipelines into PySpark frameworks.
Optimize PySpark jobs for performance, scalability, and resource utilization.
Support cloud modernization initiatives on AWS/Databricks/Snowflake platforms.
Implement ETL/ELT processes for structured and unstructured data.