Job Summary
We are looking for an experienced Python + PySpark Developer with 58 years of experience in designing, developing, and optimizing large-scale data processing solutions. The ideal candidate should have strong expertise in Python, PySpark, SQL, and big data technologies, along with experience in building scalable ETL pipelines.
Key Responsibilities
- Design, develop, and maintain scalable ETL/data processing pipelines using Python and PySpark.
- Develop and optimize Spark applications for large-scale data processing.
- Work with structured and unstructured data from multiple sources.
- Write efficient SQL queries and optimize database performance.
- Collaborate with cross-functional teams to gather and implement data requirements.
- Monitor, troubleshoot, and optimize production data pipelines.
- Ensure data quality, security,
and governance standards are followed.
- Participate in code reviews and follow best practices for software development.
Required Skills
- 5–8 years of IT experience with strong hands-on expertise in Python.
- Minimum 3+ years of experience in PySpark and Apache Spark.
- Robust proficiency in SQL and database concepts.
- Experience in developing ETL/ELT pipelines.
- Hands-on experience with Hive, Hadoop, HDFS, or similar big data technologies.
- Experience with cloud platforms such as AWS, Azure, or GCP.
- Familiarity with Git, CI/CD pipelines, and Agile methodologies.
- Strong debugging, analytical, and problem-solving skills.