Hadoop / PySpark (Bengaluru)

Hadoop / PySpark (Bengaluru)

30 Sep
|
Infosys
|
Bengaluru

30 Sep

Infosys

Bengaluru

:

- Join a data driven team where your work helps turn large scale information into clear actionable insights
- In this role you ll collaborate with engineers analysts and stakeholders to build and optimize reliable big data solutions that power reporting analytics and downstream applications
- You ll get hands on exposure to modern distributed processing contribute to production grade pipelines and learn best practices for performance quality and governance
- If you enjoy solving complex data challenges working in an environment that values curiosity and teamwork and delivering measurable impact through scalable engineering this opportunity will help you grow your technical depth while making a real difference across projects and business outcomes

Key Responsibilities:

- Design develop and support scalable data processing workflows using Hadoop and PySpark for batch and large volume processing
- Build and maintain data pipelines that ingest transform and validate data from multiple sources into curated datasets
- Optimize Spark jobs for performance partitioning caching shuffle tuning and improve overall pipeline efficiency and reliability
- Perform data quality checks reconciliation and root cause analysis for pipeline failures or data anomalies
- Collaborate with cross functional teams to understand requirements translate them into technical solutions and deliver within timelines
- Create clear technical documentation for workflows data mappings and operational runbooks




- Participate in code reviews follow engineering best practices and contribute to continuous improvement of standards and tooling

Technical Requirements:

- echnology Big Data Data Processing PySpark Technology Big Data Hadoop Hadoop Administration Hadoop

Additional Responsibilities:

- 2 5 years of experience in big data engineering or data processing roles
- Bachelor s Master s degree in Engineering Computer Science or equivalent BTech BE MCA MSc MTech or related
- Hands on experience with Hadoop ecosystem concepts HDFS distributed processing fundamentals
- Practical experience developing data transformations using PySpark
- Strong problem solving skills with the ability to debug data and job execution issues in distributed environments
- Preferred Qualifications
- Experience building end to end ETL ELT pipelines and managing dependencies across multiple data workflows
- Working knowledge of Spark optimization techniques and handling skewed large datasets efficiently
- Familiarity with data modeling concepts and designing curated datasets for analytics and reporting use cases
- Exposure to production support practices such as monitoring incident triage and improving pipeline resiliency
- Strong communication skills to collaborate effectively with stakeholders and explain technical trade offs clearly
- Valuable to have skills
- Hive HBase Kafka Airflow Scala

Preferred Skills:

Technology->Big Data - Hadoop->Hadoop Administration->Hadoop,Technology->Big Data - Data Processing->PySpark

📌 Hadoop / PySpark (Bengaluru)
🏢 Infosys
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: hadoop / pyspark (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: hadoop / pyspark (bengaluru) / bengaluru