28 Sep
|
Infosys
|
Bengaluru
- Primary skills:Technology->Big Data - Data Processing->PySpark
Key Responsibilities
- Design, develop, and maintain scalable batch data pipelines using PySpark and Apache Spark for large datasets.
- Perform data ingestion, transformation, and enrichment while ensuring accuracy, completeness, and consistency of outputs.
- Optimize Spark jobs for performance (partitioning, caching, joins, shuffles) and improve runtime efficiency and resource utilization.
- Implement robust error handling, logging, and monitoring to ensure reliable pipeline execution and faster issue resolution.
- Collaborate with cross-functional teams to gather requirements, define data contracts, and deliver well-documented solutions.
- Conduct code reviews, follow engineering best practices, and contribute to reusable components and standards.
- Troubleshoot production issues, perform root-cause analysis, and drive corrective and preventive actions.
Minimum
Qualifications:
- Bachelor’s degree (or equivalent) in Engineering/Computer Science/IT or related field.
- 3–5 years of experience in data engineering or big data development roles.
- Solid hands-on experience with PySpark and Apache Spark for building data processing workflows.
- Solid understanding of distributed data processing concepts and performance tuning fundamentals.
- Ability to translate business requirements into technical implementations and deliver within timelines.
Preferred
Qualifications:
- Experience building end-to-end Spark applications including job orchestration, dependency management, and production support readiness.
- Strong data transformation skills with a focus on data quality checks, reconciliation, and pipeline reliability.
- Exposure to designing modular, reusable Spark components and maintaining clean, maintainable codebases.
- Familiarity with structured and semi-structured data formats and efficient processing patterns in Spark.
- Proven ability to collaborate effectively across teams, communicate clearly, and contribute to continuous improvement initiatives.
Good to have skills: Spark SQL, Delta Lake, Databricks, Airflow, Hadoop
📌 Pyspark (Bengaluru)
🏢 Infosys
📍 Bengaluru