We are looking for a skilled PySpark Developer with 3 5 years of experience in designing, developing, and optimizing large-scale data processing solutions. The ideal candidate should have strong expertise in PySpark, Python, SQL, and distributed data processing frameworks. The role involves building scalable ETL pipelines, optimizing data workflows, and collaborating with cross-functional teams to deliver high-quality data solutions.
JOIN OUR TEAM
Relevant Experience
- 3 5 years
Location
- Gurugram (Work From Office)
Employment Type
- Full-time
Key Responsibilties
- Design, develop, and maintain scalable data pipelines using PySpark .
- Develop and optimize ETL/ELT workflows for processing large datasets.
- Write efficient and optimized Python and SQL code.
- Work with structured and unstructured data from multiple data sources.
- Optimize Spark jobs for performance, scalability, and reliability.
- Collaborate with Data Engineers, Data Analysts, and Business stakeholders to understand data requirements.
- Perform data validation, troubleshooting,
and root cause analysis.
- Ensure adherence to coding standards, best practices, and documentation.
- Participate in code reviews and contribute to continuous process improvements.
Required Skills
- 3 5 years of experience in Data Engineering or Big Data Development.
- Strong hands-on experience with PySpark.
- Proficiency in Python and SQL.
- Good understanding of Apache Spark architecture and optimization techniques.
- Experience in developing ETL pipelines.
- Knowledge of distributed data processing concepts.
- Experience with Linux/Unix environments.
- Robust analytical and problem-solving skills.
- Good communication and teamwork abilities.
Preferred Skills
- Experience with cloud platforms such as AWS, Azure, or GCP.
- Exposure to Databricks.
- Experience with Apache Airflow or similar workflow orchestration tools.
- Knowledge of Hive, Hadoop, or Kafka.
- Familiarity with Git and CI/CD practi