- Robust hands-on experience in PySpark
- Expertise in Apache Spark and Big Data processing
- Experience with Python programming
- Good understanding of Data Engineering concepts
- Exposure to Hive, Hadoop, SQL
- Knowledge of cloud platforms (Azure/AWS) is a plus
- Experience in ETL pipeline development and optimization
Responsibilities:
- Design, develop, and optimize PySpark-based data pipelines
- Process large-scale datasets efficiently
- Collaborate with data architects, analysts, and business teams
- Ensure data quality, performance, and scalability
- Troubleshoot and enhance existing Spark applications