- Step into a high impact consulting role where you ll lead the design and delivery of scalable big data solutions that turn complex datasets into actionable insights
- As a Lead Consultant you ll collaborate closely with data engineers architects analysts and business stakeholders to shape modern data platforms built on Hadoop ecosystems and PySpark based processing
- You ll guide teams through best practices performance tuning and reliable delivery balancing hands on problem solving with technical leadership
- This role is ideal for someone who enjoys mentoring driving technical decisions and building data pipelines that are resilient efficient and production ready
- If you re motivated by solving large scale data challenges and enabling teams to deliver measurable outcomes you ll find a collaborative environment here that values ownership clarity and continuous improvement
Key Responsibilities:
- Lead end to end delivery of big data solutions using Hadoop and PySpark from requirements to production rollout
- Design and implement scalable batch ETL pipelines for large datasets with strong focus on reliability and performance
- Drive technical architecture discussions define standards and ensure best practices for distributed data processing
- Optimize Spark jobs through partitioning strategies caching shuffle tuning and efficient file formats where applicable
- Collaborate with stakeholders to translate business needs into technical specifications delivery plans and milestones
- Establish data quality checks validation frameworks and operational monitoring for production pipelines
- Perform root cause analysis for pipeline failures and performance bottlenecks implement preventive fixes
- Mentor engineers through code reviews design reviews and knowledge sharing to uplift team capability
- Ensure documentation runbooks and handover artifacts are created and maintained for support readiness
Technical Requirements:
- Primary skills Technology Big Data Data Processing PySpark Technology Big Data Hadoop Hadoop
Additional Responsibilities:
- Bachelor s or Master s degree or equivalent in Engineering Technology Computer Applications Science BE BTech MTech MCA MSc or equivalent
- 9 11 years of overall experience in data engineering and large scale data processing environments
- Strong hands on experience with Hadoop ecosystem components and distributed data processing concepts
- Solid hands on experience building data pipelines using PySpark
- Proven ability to lead technical delivery guide teams and manage stakeholder expectations in consulting engagements
- Preferred Qualifications
- Experience designing and implementing Spark based data processing patterns batch and incremental loads
- Strong understanding of data modeling and storage patterns for big data platforms partitioning compaction schema evolution
- Experience with workflow orchestration and scheduling for data pipelines and dependency management
- Demonstrated expertise in production hardening monitoring alerting SLAs and incident management for data jobs
- Strong consulting mindset with ability to present solutions document designs clearly and influence technical decisions across teams
Preferred Skills:
Technology->Big Data - Data Processing->PySpark,Technology->Big Data - Hadoop->Hadoop