02 Aug
|
Nexhires
|
Hyderabad
02 Aug
Nexhires
Hyderabad
Key Responsibilities
Key Responsibilities
- Design, build, and maintain production-grade Apache Spark pipelines
- Develop scalable data ingestion frameworks integrating multiple source systems
- Process and transform large-scale datasets (500GB+) efficiently
- Optimize Spark applications using partitioning, caching, shuffle optimization, and performance tuning techniques
- Build reliable ETL/ELT pipelines using PySpark, Python, SQL, and Shell scripting
- Perform data cleansing, validation, normalization, and transformation to create curated datasets
- Work with Oracle SQL, HDFS, and different file formats for data processing
- Implement data quality checks, monitoring, and performance optimization practices
- Troubleshoot production issues and collaborate with cross-functional teams in Agile environments
Required Skills
Strong hands-on experience with PySpark, Python, and SQL
Deep understanding of Apache Spark architecture (Driver, Executors, DAG, Partitions, Shuffles, Caching)
Experience designing and optimizing Spark-based ETL/ELT pipelines
Strong knowledge of BigQuery and data processing frameworks
Experience with HDFS, Oracle SQL, and large-scale data processing
Understanding of data quality, governance, observability, and performance tuning
Qualification
Bachelors / Master’s degree with 6+ years of Data Engineering experience and solid PySpark expertise.
📌 PySpark Data Engineer_Immediate Joiners (Hyderabad)
🏢 Nexhires
📍 Hyderabad