Responsibilities :
- At least 3+ years of experience in designing and developing large scale, distributed data processing pipelines using PySpark, Hadoop and related technologies.
- Having expertise in Pyspark, Hadoop, Spark Core, Spark SQL, Batch processing and Spark Streaming
- Experience with Hadoop, HDFS, Hive and other BigData technologies.
- Familiarity with Data warehousing and ETL concepts and techniques
- UNIX shell scripting will be an added advantage in scheduling/running application jobs.
- At least 3 years of experience in Project development life cycle activities and development/maintenance projects
- Work with business stakeholders and other SMEs to understand high level business requirements.
- Work with the Solution Designers and contribute to the development of project plans by participating in the scoping and estimating of proposed project.
- Apply technical background understanding, business knowledge, system knowledge in the elicitation of Systems Requirements for projects.
- Possess good knowledge on Spark architecture and transformations using Spark and PySpark.
- Work in an Agile environment and participation in scrum daily standups, sprint planning reviews and retrospectives.
- Understand project requirements and translate them into technical solutions which meets the project quality standards
- Ability to work in team in diverse/multiple stakeholder environment and collaborate with upstream/downstream functional teams to identify, troubleshoot and resolve data issues.
Additional Responsibilities :
Infosys is a global leader in next-generation digital services and consulting. Within the Data & Analytics unit, you will work on cutting-edge data engineering and analytics initiatives for global clients across industries. The role offers opportunities to work with cloud-native data platforms, modern analytics ecosystems, AI-driven solutions, and large-scale enterprise data modernization programs. Employees benefit from continuous learning programs, certifications, and career growth opportunities within Infosys' data and analytics practice
Technical and Professional Requirements :
Primary skills:
Pyspark, Hadoop, Spark Core, Spark SQL, Batch processing and Spark Streaming
Hadoop, HDFS, Hive and other BigData technologies.
Strong problem solving and Positive Analytical skills.
- Excellent verbal and written communication skills.
- Experience and desire to work in a Global delivery environment.
- Stay up to date with new technologies and industry trends in Development.
Preferred Skills :
Technology->Big Data - Data Processing->PySpark
Educational Requirements :
Bachelor of Engineering,Bachelor Of Technology,Bachelor Of Comp. Applications,Bachelor Of Science,Master Of Technology,Master Of Comp. Applications
Service Line :
Data & Analytics Unit
📌 Pyspark Developer (Pune)
🏢 Infosys
📍 Pune