- Design, develop, test, deploy and maintain large-scale data pipelines using Delta Lake and PySpark on AWS.
- Collaborate with cross-functional teams to gather requirements and deliver high-quality solutions that meet business needs.
- Develop complex SQL queries to optimize database performance and troubleshoot issues in a distributed computing environment.
- Ensure scalability, reliability, security, and compliance of all systems by implementing best practices in software development.
Job Requirements :
- 7-12 years of experience in designing and developing big data solutions using Python (PySpark), SQL, and other technologies such as Data Bricks.
- Robust understanding of Delta Lake architecture and its applications in real-world scenarios.
- Proficiency in writing efficient SQL queries for querying large datasets using various databases like MySQL or PostgreSQL.