28 Sep
|
Infosys
|
Bengaluru
- Primary skills:Technology->Big Data - Data Processing->PySpark,Technology->Big Data - Hadoop->Hadoop
- Lead end-to-end delivery of big data solutions using Hadoop and PySpark, from requirements to production rollout.
- Design and implement scalable batch/ETL pipelines for large datasets with strong focus on reliability and performance.
- Drive technical architecture discussions, define standards, and ensure best practices for distributed data processing.
- Optimize Spark jobs through partitioning strategies, caching, shuffle tuning, and efficient file formats where applicable.
- Collaborate with stakeholders to translate business needs into technical specifications, delivery plans, and milestones.
- Establish data quality checks, validation frameworks, and operational monitoring for production pipelines.
- Perform root-cause analysis for pipeline failures and performance bottlenecks; implement preventive fixes.
- Mentor engineers through code reviews, design reviews, and knowledge sharing to uplift team capability.
- Ensure documentation, runbooks, and handover artifacts are created and maintained for support readiness.
- Bachelor’s or Master’s degree (or equivalent)
in Engineering/Technology/Computer Applications/Science (BE/BTech/MTech/MCA/MSc or equivalent).
- 9–11 years of overall experience in data engineering and large-scale data processing environments.
- Solid hands-on experience with Hadoop ecosystem components and distributed data processing concepts.
- Strong hands-on experience building data pipelines using PySpark.
- Proven ability to lead technical delivery, guide teams, and manage stakeholder expectations in consulting engagements.
Preferred
Qualifications
- Experience designing and implementing Spark-based data processing patterns (batch and incremental loads).
- Strong understanding of data modeling and storage patterns for big data platforms (partitioning, compaction, schema evolution).
- Experience with workflow orchestration and scheduling for data pipelines and dependency management.
- Demonstrated expertise in production hardening: monitoring, alerting, SLAs, and incident management for data jobs.
- Strong consulting mindset with ability to present solutions, document designs clearly, and influence technical decisions across teams.
📌 Hadoop / PySpark (Bengaluru)
🏢 Infosys
📍 Bengaluru