- Develop and maintain large-scale data pipelines using Cloudera Hadoop ecosystem.
- Work hands-on with Cloudera 7.1+ in the current project.
- Develop data processing solutions using PySpark, Hive, and SQL.
- Work with Apache Iceberg for data lake/lakehouse environments.
- Troubleshoot and optimize Spark jobs and SQL queries.
- Support data integration, transformation, and production data pipelines.
- Work with ODCS (Open Data Contract Standard) where applicable.