- Strong programming experience in Python
- Hands-on expertise in PySpark
- Experience with AWS services such as:
- S3
- EMR
- Glue
- Lambda
- IAM
- Athena
- Redshift
- Robust SQL and Data Warehousing concepts
- Experience in building ETL/ELT pipelines
- Data Lake/Data Warehouse implementation experience
- Git/GitHub version control
- Performance tuning and optimization of Spark applications
- Understanding of Agile development methodologies