- Join a fast paced collaborative team where data powers smarter decisions and better customer experiences
- In this role you ll work hands on with large scale datasets to build reliable high performing data processing solutions using PySpark and Spark
- You ll partner closely with engineers analysts and stakeholders to understand business needs translate them into scalable pipelines and continuously improve data quality and performance
- If you enjoy solving complex data challenges optimizing distributed workloads and taking ownership from design to delivery this is a great opportunity to grow your impact
- You ll be encouraged to share ideas learn from peers and contribute to a culture that values clarity craftsmanship and continuous improvement
Key Responsibilities:
- Key Responsibilities
- Design develop and maintain scalable batch data pipelines using PySpark and Apache Spark for large datasets
- Perform data ingestion transformation and enrichment while ensuring accuracy completeness and consistency of outputs
- Optimize Spark jobs for performance partitioning caching joins shuffles and improve runtime efficiency and resource utilization
- Implement robust error handling logging and monitoring to ensure reliable pipeline execution and faster issue resolution
- Collaborate with cross functional teams to gather requirements define data contracts and deliver well documented solutions
- Conduct code reviews follow engineering best practices and contribute to reusable components and standards
- Troubleshoot production issues perform root cause analysis and drive corrective and preventive actions
Technical Requirements:
- Primary skills Technology Big Data Data Processing PySpark
Additional Responsibilities:
- Minimum Qualifications
- Bachelor s degree or equivalent in Engineering Computer Science IT or related field
- 3 5 years of experience in data engineering or big data development roles
- Strong hands on experience with PySpark and Apache Spark for building data processing workflows
- Solid understanding of distributed data processing concepts and performance tuning fundamentals
- Ability to translate business requirements into technical implementations and deliver within timelines
- Preferred Qualifications
- Experience building end to end Spark applications including job orchestration dependency management and production support readiness
- Strong data transformation skills with a focus on data quality checks reconciliation and pipeline reliability
- Exposure to designing modular reusable Spark components and maintaining clean maintainable codebases
- Familiarity with structured and semi structured data formats and effective processing patterns in Spark
- Proven ability to collaborate effectively across teams communicate clearly and contribute to continuous improvement initiatives
- Good to have skills
- Spark SQL Delta Lake Databricks Airflow Hadoop
Preferred Skills:
Technology->Big Data - Data Processing->PySpark
📌 Pyspark (Bengaluru)
🏢 Infosys
📍 Bengaluru
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.