- Design, develop, and deploy large-scale data processing pipelines using PySpark and Spark.
- Collaborate with cross-functional teams to identify business requirements and design solutions.
- Develop and maintain high-quality code that meets performance, security, and reliability standards.
- Troubleshoot and resolve complex technical issues related to data processing and pipeline failures.
- Participate in code reviews and contribute to the improvement of the overall codebase quality.
- Stay up-to-date with industry trends and emerging technologies in big data processing.
Job Requirements
- Strong proficiency in Python programming language and its ecosystem.
- Experience with PySpark and Spark is required; knowledge of other big data frameworks like Hadoop or Flink is a plus.
- Excellent problem-solving skills and attention to detail.
- Ability to work collaboratively in a team workplace and communicate effectively with stakeholders.
- Strong understanding of software development principles, patterns, and practices.
- Familiarity with agile development methodologies and version control systems like Git.