- Design, develop, test, deploy, and maintain large-scale data pipelines using PySpark on AWS EMR, Azure Data Factory, or similar platforms.
- Collaborate with cross-functional teams to gather requirements and deliver high-quality solutions that meet business needs.
- Develop complex SQL queries to extract insights from massive datasets stored in MySQL databases.
- Implement machine learning models using Python libraries such as scikit-learn or TensorFlow.
Desired Candidate Profile
- 2-5 years of experience in developing scalable big data applications using Hadoop ecosystem (Hive) and Spark (PySpark).
- Robust proficiency in programming languages like Scala/Python/Java; familiarity with SQL Server/SQL Azure/MySQL preferred.
- Experience working with cloud-based technologies like AWS/Azure Data Factory/Microsoft Azure; knowledge of machine learning algorithms an added advantage.
- Bachelor's degree in Any Specialization (B.C.A./B.Sc.).