30 Jul
|
Nexhires
|
Telangana
30 Jul
Nexhires
Telangana
Key Responsibilities
Key Responsibilities:
• Design, build, and maintain production-grade Apache Spark pipelines
• Develop scalable data ingestion frameworks integrating multiple source systems
• Process and transform large-scale datasets (500GB+) efficiently
• Optimize Spark applications using partitioning, caching, shuffle optimization, and performance tuning techniques
• Build reliable ETL/ELT pipelines using PySpark, Python, SQL, and Shell scripting
• Perform data cleansing, validation, normalization, and transformation to create curated datasets
• Work with Oracle SQL, HDFS, and different file formats for data processing
• Implement data quality checks, monitoring, and performance optimization practices
• Troubleshoot production issues and collaborate with cross-functional teams in Agile environments
Required Skills:
Strong hands-on experience with PySpark, Python, and SQL
Deep understanding of Apache Spark architecture (Driver, Executors, DAG, Partitions, Shuffles, Caching)
Experience designing and optimizing Spark-based ETL/ELT pipelines
Robust knowledge of BigQuery and data processing frameworks
Experience with HDFS, Oracle SQL, and large-scale data processing
Understanding of data quality, governance, observability, and performance tuning
Qualification:
Bachelors / Master’s degree with 6+ years of Data Engineering experience and strong PySpark expertise.
📌 PySpark Data Engineer_Immediate Joiners (Telangana)
🏢 Nexhires
📍 Telangana