We are looking for an experienced Python + PySpark Developer with 58 years of experience in designing, developing, and optimizing large-scale data processing solutions. The ideal candidate should have solid expertise in Python, PySpark, SQL, and big data technologies, along with experience in building scalable ETL pipelines.
Key Responsibilities
Design, develop, and maintain scalable ETL/data processing pipelines using Python and PySpark.
Develop and optimize Spark applications for large-scale data processing.
Work with structured and unstructured data from multiple sources.
Write efficient SQL queries and optimize database performance.
Collaborate with cross-functional teams to gather and implement data requirements.
Monitor, troubleshoot, and optimize production data pipelines.
Ensure data quality, security,
and governance standards are followed.
Participate in code reviews and follow best practices for software development.
Required Skills
5–8 years of IT experience with strong hands-on expertise in Python.
Minimum 3+ years of experience in PySpark and Apache Spark.
Solid proficiency in SQL and database concepts.
Experience in developing ETL/ELT pipelines.
Hands-on experience with Hive, Hadoop, HDFS, or similar big data technologies.
Experience with cloud platforms such as AWS, Azure, or GCP.
Familiarity with Git, CI/CD pipelines, and Agile methodologies.
Solid debugging, analytical, and problem-solving skills.