Lead Data Engineer (Hadoop)
Location: Hyderabad
Experience: 6 to 10 Years
Employment Type: Full-Time
Immediate joiners
Key Responsibilities
- Design, build, and maintain scalable ETL/ELT data pipelines using Hadoop, Hive, HDFS, and Spark/PySpark
- Develop and optimize complex SQL queries and Hive scripts for large-scale data processing
- Write clean, efficient, and reusable Python code for data transformation and automation
- Work with structured and unstructured data across distributed storage systems (HDFS)
- Optimize Spark/PySpark jobs for performance, scalability, and resource efficiency
- Collaborate with data analysts, data scientists, and business stakeholders to understand data requirements
- Ensure data quality, consistency, and integrity across pipelines
- Troubleshoot and resolve issues related to data pipeline failures, performance bottlenecks, and cluster resource management
- Participate in code reviews and follow best practices for data engineering and version control
- Document technical designs, data flows, and pipeline architecture
Required Skills & Experience
- 6 to 10 years of hands-on experience in Data Engineering
- Strong working knowledge of Hadoop ecosystem (HDFS, YARN, MapReduce concepts)
- Proficiency in Hive for data warehousing and query optimization
- Solid experience with Spark/PySpark for distributed data processing
- Strong programming skills in Python
- Advanced SQL skills — query optimization, joins, window functions, performance tuning
Positive to Have
- Experience with NoSQL databases (HBase, Cassandra)
- Familiarity with CI/CD pipelines for data engineering workflows
Educational Qualification
- Bachelor's or Master's degree in Computer Science, Information Technology, or a related field
Regards,
Manvendra Singh
[email protected]
📌 Lead Data Engineer - Hadoop (Hyderabad)
🏢 Incedo
📍 Hyderabad