We are seeking a skilled PySpark Developer with 2-8 years of experience in designing, developing, and optimizing large-scale data processing solutions. The ideal candidate should have hands-on experience with PySpark, Python, Spark SQL, ETL development, and Big Data technologies. The role involves building scalable data pipelines, transforming large datasets, and supporting analytics and reporting requirements.
Key Responsibilities
- Design, develop, and maintain scalable data processing applications using PySpark.
- Build and optimize ETL/ELT pipelines for structured and unstructured data.
- Develop data transformation logic using Spark SQL and DataFrames.
- Work with large-scale datasets in distributed computing environments.
- Collaborate with Data Engineers, Data Analysts, and Business stakeholders to understand data requirements.
- Monitor, troubleshoot, and optimize Spark jobs for performance and reliability.
- Ensure data quality, integrity, and consistency across data pipelines.
- Participate in code reviews and follow coding best practices.
- Support deployments, enhancements, and production issue resolution.
- Create technical documentation and maintain operational procedures.