28 Sep
|
Infosys
|
Bengaluru
Good to have skills: Spark SQL, YARN, HDFS, Oozie, Airflow
Key Responsibilities
- Lead the design and development of scalable Big Data solutions using Hadoop and PySpark for batch and large-scale processing.
- Architect and implement end-to-end data pipelines, ensuring reliability, performance tuning, and efficient resource utilization on Hadoop clusters.
- Develop and optimize Hive data models, queries, and partitioning strategies to support analytics and downstream consumption.
- Drive technical planning, estimation, and delivery for data engineering initiatives, ensuring timelines and quality standards are met.
- Establish coding standards, review code, and enforce best practices for maintainability, testing, and production readiness.
- Troubleshoot production issues, perform root-cause analysis, and implement preventive measures to improve stability and throughput.
- Collaborate with product, analytics, and platform teams to translate requirements into scalable technical solutions.
- Mentor team members, guide technical decisions, and support skill development across Hadoop, PySpark, Big Data, and Hive.
Minimum
Qualifications:
- Education: BTECH, MTECH, MCA, MSC (or equivalent).
- 5–9 years of overall experience with strong hands-on expertise in Hadoop and PySpark for large-scale data processing.
- Proven experience building and maintaining Big Data pipelines and working with Hive for querying and data modeling.
- Solid understanding of distributed processing concepts, performance optimization, and data reliability practices.
- Experience leading technical execution through code reviews, design discussions, and delivery ownership.
Preferred
Qualifications:
- Experience designing reusable frameworks and standardized pipeline patterns to improve team productivity and consistency.
- Strong expertise in optimizing Spark jobs (partitioning, caching, shuffles) and Hive performance (file formats, partitions, bucketing).
- Experience implementing data quality checks, monitoring, and operational dashboards for production pipelines.
- Ability to drive stakeholder communication, manage technical trade-offs, and lead solutioning for complex data use cases.
- Demonstrated mentoring and leadership experience, enabling teams to deliver high-quality Big Data solutions at scale.
📌 Hadoop / PySpark (Bengaluru)
🏢 Infosys
📍 Bengaluru