Roles and Responsibilities
- Design, develop, test, deploy, and maintain large-scale data processing pipelines using PySpark on AWS.
- Collaborate with cross-functional teams to gather requirements and deliver high-quality solutions.
- Develop complex ETL processes to extract insights from structured and unstructured data sources.
- Ensure scalability, performance, and reliability of the developed applications.
Desired Candidate Profile
- 2-5 years of experience in PySpark development with expertise in AWS services such as S3, Glue, Lambda etc. .
- Robust understanding of SQL concepts including joins, subqueries, aggregations etc. .
- Experience working with Python programming language with knowledge of its libraries like NumPy pandas etc. .
- Bachelor's degree in Any Specialization (B.C.A. / B.Sc.).
- Hands-on experience with DataBricks framework for building scalable data pipelines.
📌 Pyspark (Data Engineer)-Pan India-2-5yrs (Pune)
🏢 Infosys
📍 Pune