Design, develop, and maintain scalable data pipelines using Python and PySpark
Develop and optimize ELT/ETL data processing solutions
Build and support batch and real-time/streaming data processing solutions
Work with large-scale datasets in distributed environments
Develop, modify, and troubleshoot UNIX/Linux Shell scripts
Write and optimize SQL queries for data processing and transformation
Perform PySpark job performance tuning and troubleshooting
Design and implement scalable data engineering solutions
Gather requirements and contribute to solution design and technical documentation
Support and troubleshoot data engineering applications and workflows
Collaborate effectively with technical and business stakeholders
Required Skills:
Strong hands-on experience in Python programming, coding, data processing, and performance optimization
Extensive hands-on experience in PySpark development
Strong experience in PySpark application design, job execution, performance tuning, and troubleshooting
Hands-on experience in SQL coding
Experience in ELT/ETL pipeline development and data integration
Proficiency in UNIX/Linux Shell Scripting
Experience with batch and real-time/streaming data processing
Ability to handle large-scale data processing and optimize workflows for performance and scalability
Valuable understanding of distributed computing and Big Data architecture
Experience in requirement gathering, solution design, technical documentation, and implementation
Strong analytical and problem-solving skills
Excellent verbal and written communication skills
📌 Sr. Pyspark developer (Bengaluru)
🏢 EduRun Group
📍 Bengaluru
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.