Description
Job Title: Data Engineer (5+ Years Experience)
Job Summary
We are seeking a Data Engineer with 5+ years of experience in building and maintaining data pipelines using PySpark, SQL, and Python. The candidate should have a solid understanding of Bigdata tools like Hadoop, Hive, Oozie and have basic knowledge of cloud environments (preferably AWS). The role requires working closely with teams to support data processing and analytics needs.
Key Responsibilities
Develop and maintain data pipelines using PySpark and Python, Hive, Oozie
Write productive SQL queries for data extraction, transformation, and validation
Assist in integrating data from multiple sources and ensuring data accuracy
Support debugging, monitoring, and optimization of data pipelines
Collaborate with team members to understand data requirements and deliver solutions
Follow best practices for data engineering and documentation
Required Skills (Primary)
5+ years of hands-on experience in PySpark, SQL, and Python
5+ years of hands-on experience in Big Data Tools like Hive ,Oozie
Working knowledge of cloud environments (preferably AWS)
Understanding of data processing and ETL concepts
Secondary Skills
Basic experience with Databricks
Familiarity with Linux/Unix commands for working on edge nodes
Exposure to scheduling or orchestration tools
Good to Have
Basic understanding of finance domain
Exposure to Hive and Oozie
Understanding of data warehousing concepts
DBT, DAGSTER
Qualifications
Bachelor's degree in computer science, Engineering, or related field
5+ years of relevant experience in data engineering or big data technologies
Soft Skills
Good analytical and problem-solving skills
Effective communication and teamwork abilities
Willingness to learn and adapt in a quick-paced setting
Responsibilities
Develop and maintain data pipelines using PySpark and Python, Hive, Oozie
Write efficient SQL queries for data extraction, transformation, and validation
Assist in integrating data from multiple sources and ensuring data accuracy
Support debugging, monitoring, and optimization of data pipelines
Collaborate with team members to understand data requirements and deliver solutions
Follow best practices for data engineering and documentation
Qualifications
Graduate in Computer Science, Data Science, or related field. 5+ years of experience in data engineering or related field.
📌 Assistant Manager Pune (India)
🏢 EXL
📍 India