28 Sep
|
Infosys
|
Bangalore East
28 Sep
Infosys
Bangalore East
Valuable to have skills: SQL, Hadoop, Hive, Kafka, Airflow
Key Responsibilities
- Design, develop, and maintain scalable ETL/ELT pipelines using PySpark for batch and/or incremental processing.
- Build and optimize Apache Spark jobs with focus on performance, partitioning strategy, caching, and efficient transformations/actions.
- Perform data cleansing, validation, and reconciliation to ensure accuracy, completeness, and consistency of datasets.
- Collaborate with cross-functional teams to understand requirements and translate them into robust data processing solutions.
- Troubleshoot pipeline failures, analyze logs, identify bottlenecks, and implement fixes to improve reliability and throughput.
- Write clean, maintainable code with reusable components and clear documentation for pipelines and data flows.
- Support deployment and operationalization of Spark workloads, including monitoring and basic production support activities.
- Contribute to code reviews and follow engineering best practices to improve quality and maintainability.
Minimum
Qualifications:
- Education: BTECH,
MTECH, MCA, MSC.
- 2–3 years of experience in data engineering or large-scale data processing roles.
- Strong hands-on experience with PySpark for building data pipelines and transformations.
- Working knowledge of Apache Spark concepts such as RDD/DataFrame, joins, shuffles, and performance considerations.
- Ability to debug Spark applications and resolve data/job issues effectively.
Preferred
Qualifications:
- Experience optimizing Spark workloads (tuning partitions, managing skew, memory/executor settings) for performance and cost efficiency.
- Exposure to building end-to-end data pipelines with strong data quality checks and automated validations.
- Familiarity with distributed processing patterns and designing reusable PySpark modules for scalable development.
- Experience collaborating in agile teams, participating in code reviews, and improving engineering standards for data pipelines.
📌 Pyspark (Bangalore East)
🏢 Infosys
📍 Bangalore East