18 Aug
|
Objectwin Technology India
|
India
18 Aug
Objectwin Technology India
India
Role Summary
We are looking for an experienced PySpark Developer with strong hands-on expertise in big data processing, distributed computing, and data engineering. The ideal candidate will have deep experience building scalable data pipelines, transforming large datasets, and working with Spark-based ecosystems in production environments.
Key Responsibilities
- Design, build, and maintain scalable data pipelines using PySpark and Apache Spark
- Develop efficient ETL/ELT workflows for batch and near-real-time processing
- Optimize Spark jobs for performance, reliability, and cost efficiency
- Work with large structured and unstructured datasets
- Integrate data from multiple sources such as databases, APIs, files, and cloud storage
- Write reusable, modular, and maintainable PySpark code
- Troubleshoot job failures, data quality issues, and performance bottlenecks
- Collaborate with data architects, analysts, platform teams, and business stakeholders
- Implement data validation, monitoring, and logging frameworks
- Support deployment, scheduling, and orchestration of data pipelines
- Participate in design reviews, code reviews, and technical discussions
- Mentor junior engineers and contribute to team best practices
Required Skills
- Solid hands-on experience with PySpark and Apache Spark
- Deep understanding of Spark concepts such as RDDs, Data Frames, datasets, partitioning, caching, shuffling, joins, and window functions
- Strong Python programming skills
- Experience with SQL and relational databases
- Knowledge of big data concepts and distributed data processing
- Hands-on experience with ETL/ELT pipeline development
- Good understanding of performance tuning and optimization techniques in Spark
- Experience with version control tools like Git
- Familiarity with Linux/Unix environments
- Strong debugging and analytical skills
📌 PySpark Developer / Senior Data Engineer 6+ (India)
🏢 Objectwin Technology India
📍 India