- Design, develop, test, deploy and maintain large-scale data processing pipelines using PySpark.
- Collaborate with cross-functional teams to gather requirements and deliver high-quality solutions on time.
- Develop complex SQL queries to extract insights from massive datasets stored in relational databases.
- Troubleshoot issues related to data quality, performance optimization, and scalability.
Job Requirements :
- 3-6 years of experience in developing applications using Python programming language.
- Solid understanding of SQL concepts and ability to write efficient queries for large datasets.
- Proficiency in working with PySpark framework for big data processing tasks.
- Experience with database management systems such as MySQL or PostgreSQL.