Pyspark (Bengaluru)

Pyspark (Bengaluru)

30 Sep
|
Infosys
|
Bengaluru

30 Sep

Infosys

Bengaluru

:

- About the job
- Join a collaborative data engineering team where your work directly powers reliable analytics and smarter business decisions
- In this role you ll build and optimize scalable data processing pipelines using PySpark and Spark working closely with engineers analysts and stakeholders to turn raw data into trusted high quality datasets
- You ll be encouraged to take ownership suggest improvements and contribute to a culture that values clean engineering performance and continuous learning
- If you enjoy solving data challenges tuning distributed jobs and delivering dependable solutions in a fast moving environment this is a excellent opportunity to grow your impact while working with modern big data technologies

Key Responsibilities:

- Key Responsibilities
- Design develop and maintain scalable ETL ELT pipelines using PySpark for batch and or incremental processing
- Build and optimize Apache Spark jobs with focus on performance partitioning strategy caching and efficient transformations actions
- Perform data cleansing validation and reconciliation to ensure accuracy completeness and consistency of datasets
- Collaborate with cross functional teams to understand requirements and translate them into robust data processing solutions
- Troubleshoot pipeline failures analyze logs identify bottlenecks and implement fixes to improve reliability and throughput




- Write clean maintainable code with reusable components and clear documentation for pipelines and data flows
- Support deployment and operationalization of Spark workloads including monitoring and basic production support activities
- Contribute to code reviews and follow engineering best practices to improve quality and maintainability
- Minimum Qualifications
- Education BTECH MTECH MCA MSC
- 2 3 years of experience in data engineering or large scale data processing roles
- Strong hands on experience with PySpark for building data pipelines and transformations
- Working knowledge of Apache Spark concepts such as RDD DataFrame joins shuffles and performance considerations
- Ability to debug Spark applications and resolve data job issues effectively

Technical Requirements:

- Good to have skills
- SQL Hadoop Hive Kafka Airflow

Additional Responsibilities:

- Preferred Qualifications
- Experience optimizing Spark workloads tuning partitions managing skew memory executor settings for performance and cost efficiency
- Exposure to building end to end data pipelines with strong data quality checks and automated validations
- Familiarity with distributed processing patterns and designing reusable PySpark modules for scalable development
- Experience collaborating in agile teams participating in code reviews and improving engineering standards for data pipelines

Preferred Skills:

Technology->Big Data - Data Processing->PySpark

📌 Pyspark (Bengaluru)
🏢 Infosys
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: pyspark (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: pyspark (bengaluru) / bengaluru