- About the job
- Step into a leadership role where you ll shape modern data platforms and help teams turn large scale data into reliable high impact insights
- As a technology lead you ll work at the intersection of engineering excellence and collaboration guiding design decisions setting delivery standards and enabling your team to build scalable solutions using Hadoop and PySpark
- You ll partner closely with stakeholders to understand business needs translate them into robust data pipelines and ensure performance quality and governance across the ecosystem
- If you enjoy solving complex data challenges mentoring engineers and driving best practices in Big Data processing this role offers the opportunity to lead meaningful work while building a culture of ownership learning and continuous improvement
Key Responsibilities:
- Key Responsibilities
- Lead the design and development of scalable Big Data solutions using Hadoop and PySpark for batch and large scale processing
- Architect and implement end to end data pipelines ensuring reliability performance tuning and efficient resource utilization on Hadoop clusters
- Develop and optimize Hive data models queries and partitioning strategies to support analytics and downstream consumption
- Drive technical planning estimation and delivery for data engineering initiatives ensuring timelines and quality standards are met
- Establish coding standards review code and enforce best practices for maintainability testing and production readiness
- Troubleshoot production issues perform root cause analysis and implement preventive measures to improve stability and throughput
- Collaborate with product analytics and platform teams to translate requirements into scalable technical solutions
- Mentor team members guide technical decisions and support skill development across Hadoop PySpark Big Data and Hive
- Minimum Qualifications
- Education BTECH MTECH MCA MSC or equivalent
- 5 9 years of overall experience with strong hands on expertise in Hadoop and PySpark for large scale data processing
- Proven experience building and maintaining Big Data pipelines and working with Hive for querying and data modeling
- Robust understanding of distributed processing concepts performance optimization and data reliability practices
- Experience leading technical execution through code reviews design discussions and delivery ownership
Technical Requirements:
- Good to have skills
- Spark SQL YARN HDFS Oozie Airflow
Additional Responsibilities:
- Preferred Qualifications
- Experience designing reusable frameworks and standardized pipeline patterns to improve team productivity and consistency
- Strong expertise in optimizing Spark jobs partitioning caching shuffles and Hive performance file formats partitions bucketing
- Experience implementing data quality checks monitoring and operational dashboards for production pipelines
- Ability to drive stakeholder communication manage technical trade offs and lead solutioning for complex data use cases
- Demonstrated mentoring and leadership experience enabling teams to deliver high quality Big Data solutions at scale
Preferred Skills:
Technology->Big Data - Hadoop->Hadoop,Technology->Big Data - Data Processing->PySpark