- Design, develop, test, and deploy scalable data pipelines using Python and PySpark.
- Collaborate with cross-functional teams to gather requirements and deliver high-quality solutions on time.
- Develop complex algorithms for data processing, analysis, and visualization using advanced Python libraries such as NumPy, Pandas, Matplotlib.
- Troubleshoot issues related to data quality, performance optimization, and system integration.
Job Requirements :
- 7-9 years of experience in Python development with expertise in PySpark.
- Robust understanding of big data technologies including Hadoop ecosystem (HDFS, YARN) and Spark (Scala).
- Proficiency in developing scalable applications using Python frameworks like Django or Flask.
- Experience working with cloud-based platforms such as AWS or GCP.