We are seeking a skilled
PySpark Developer
to join our Data Engineering team. The ideal candidate should possess strong hands-on experience in
PySpark, Python, and SQL
, with expertise in developing scalable data processing solutions and ETL pipelines. The candidate will be responsible for designing, developing, optimizing, and maintaining high-performance data applications that support enterprise analytics and business intelligence initiatives.
Key Responsibilities
- Design, develop, and maintain scalable ETL/ELT pipelines using PySpark.
- Develop data processing solutions for large-scale structured and unstructured datasets.
- Build, optimize, and troubleshoot Spark applications for performance and scalability.
- Perform data extraction, transformation, validation, and loading activities.
- Develop reusable and efficient data engineering frameworks and components.
- Collaborate with business analysts, data architects,
and cross-functional teams to understand data requirements.
- Ensure data quality, integrity, and governance across data platforms.
- Monitor and support production data pipelines and resolve performance bottlenecks.
- Participate in code reviews and follow best practices for development and deployment.
Required Skills
- Robust hands-on experience with
PySpark
- Proficiency in
Python
- Strong knowledge of
SQL
- Experience with ETL/Data Engineering projects
- Good understanding of Data Warehousing concepts
- Experience with Spark SQL, DataFrames, and distributed data processing
- Knowledge of performance tuning and optimization techniques in Spark
- Experience working in Linux/Unix environments
- Strong analytical and problem-solving skills