09 Sep
|
EduRun Group
|
Bengaluru
09 Sep
EduRun Group
Bengaluru
- Design, develop, and maintain scalable data pipelines using Python and PySpark
- Develop and optimize ELT/ETL data processing solutions
- Build and support batch and real-time/streaming data processing solutions
- Work with large-scale datasets in distributed environments
- Develop, modify, and troubleshoot UNIX/Linux Shell scripts
- Write and optimize SQL queries for data processing and transformation
- Perform PySpark job performance tuning and troubleshooting
- Design and implement scalable data engineering solutions
- Gather requirements and contribute to solution design and technical documentation
- Support and troubleshoot data engineering applications and workflows
- Collaborate effectively with technical and business stakeholders
Required Skills:
- Strong hands-on experience in Python programming, coding, data processing, and performance optimization
- Extensive hands-on experience in PySpark development
- Strong experience in PySpark application design, job execution, performance tuning, and troubleshooting
- Hands-on experience in SQL coding
- Experience in ELT/ETL pipeline development and data integration
- Proficiency in UNIX/Linux Shell Scripting
- Experience with batch and real-time/streaming data processing
- Ability to handle large-scale data processing and optimize workflows for performance and scalability
- Good understanding of distributed computing and Big Data architecture
- Experience in requirement gathering, solution design, technical documentation, and implementation
- Robust analytical and problem-solving skills
- Excellent verbal and written communication skills
📌 Sr. Pyspark developer (Bengaluru)
🏢 EduRun Group
📍 Bengaluru