We are looking for a skilled PySpark Developer with 2-8 years of experience in designing, developing, and optimizing large-scale data processing solutions. The ideal candidate should have strong expertise in PySpark, Python, SQL, and Big Data technologies, with hands-on experience in building scalable ETL pipelines and data engineering solutions.
Key Responsibilities
- Design, develop, and maintain scalable data processing applications using PySpark.
- Build and optimize ETL/ELT pipelines for processing large datasets.
- Work with structured and unstructured data from multiple sources.
- Develop effective data transformation and data quality frameworks.
- Collaborate with data architects, analysts, and business stakeholders to understand requirements.
- Optimize Spark jobs for performance, scalability, and reliability.
- Perform code reviews, debugging, and troubleshooting of data pipelines.
- Implement best practices for data governance, security, and compliance.
- Support production deployments and resolve data-related issues.