- Bachelor’s degree with 2-6 years of proven experience in Python
- Robust programming skills in Python.
- Hands-on experience with PySpark, Spark SQL, and distributed data processing.
- Good understanding of ETL processes, data wrangling, and performance tuning in Spark.
- Experience with SQL, data lakes, and big data tools (e.g., Hadoop, Hive – optional).
- Familiarity with version control (Git), workflow orchestration tools (e.g., Airflow), and cloud platforms (AWS/GCP/Azure – optional).
- Strong analytical thinking, problem-solving skills, and ability to work in Agile teams.