- Design, develop, test, deploy and maintain ETL processes using PySpark to extract data from various sources.
- Collaborate with cross-functional teams to identify business requirements and design solutions that meet those needs.
- Develop complex SQL queries to optimize database performance and troubleshoot issues.
- Ensure high-quality delivery of projects by adhering to coding standards, best practices, and quality assurance processes.
Job Requirements :
- 5-12 years of experience in ETL development with a focus on big data technologies like Hadoop/Hive/Pyspark.
- Solid proficiency in Python programming language with expertise in working with PySpark libraries.
- Experience with relational databases (e.g., MySQL) as well as NoSQL databases (e.g., MongoDB).
- Proven track record of delivering large-scale ETL projects on time while ensuring high-quality code.