16 Aug
|
Web Spiders
|
Kolkata
16 Aug
Web Spiders
Kolkata
Web Spiders is looking for a Senior Spark / PySpark Data Engineer with strong hands-on experience in Apache Spark, PySpark, Python, and large-scale data engineering.
The ideal candidate will have strong experience designing and developing high-performance ETL/ELT pipelines and distributed data-processing solutions using Spark/PySpark, along with experience working with cloud-based data platforms and AWS services.
If Spark + PySpark + Python is your core expertise and you enjoy solving complex data-processing and scalability challenges, we'd love to hear from you.
5+ Years Experience | Kolkata – Work from Office
Core Stack: Apache Spark
- PySpark
- Python
- ETL/ELT
- AWS
- S3
- Glue
- Redshift
- Immediate joiners preferred.*
Working Hours: Ability to work in the US Eastern Time Zone. Depending on project requirements, this may be adjusted to a half-day IST + half-day US EST schedule.
What You'll Do
- Design, develop, and optimize large-scale ETL/ELT pipelines using Apache Spark and PySpark.
- Develop scalable data transformation and processing solutions using PySpark and Python.
- Build distributed data-processing applications capable of handling large volumes of data.
- Develop reusable and maintainable Spark/PySpark frameworks and data-processing components.
- Optimize Spark jobs for performance, scalability, memory utilization, and execution efficiency.
- Work with complex transformations, joins, aggregations, partitioning, and large datasets.
- Implement data validation, quality checks, error handling, and monitoring within data pipelines.
- Work with AWS data services including EMR, Glue, S3, and Redshift.
- Develop data pipelines supporting data lakes, warehouses, analytics, and downstream applications.
- Troubleshoot production data pipeline and Spark processing issues.
- Identify and resolve performance bottlenecks in Spark/PySpark workloads.
- Collaborate with Data Engineering, Cloud, AI/ML, and Product teams to deliver reliable data solutions.
Must-Have Skills
- 5+ years of hands-on experience in Data Engineering.
- Strong hands-on experience with Apache Spark.
- Strong hands-on experience with PySpark.
- Strong programming experience in Python.
- Proven experience developing and optimizing large-scale ETL/ELT pipelines.
- Strong understanding of distributed computing and data-processing concepts.
- Experience working with large datasets and complex data transformations.
- Strong understanding of Spark performance optimization and tuning.
- Experience with cloud-based data engineering, preferably AWS.
- Experience with Amazon S3 and at least one AWS data-processing service such as EMR or Glue.
Valuable To Have
- AWS EMR
- AWS Glue
- Apache Airflow / MWAA
- AWS Step Functions
- Amazon Redshift
- Hadoop ecosystem
- Experience with data lake and data warehouse architectures.
- Experience with CI/CD and production deployment of data pipelines.
- AWS Certified Data Engineer or another relevant AWS certification.
Interview Process
- Application review
- 5–10 minute initial screening call with the TA team
- Technical interviews {Domain specific}
- Practical test conducted in the presence of a panel member
- Role match & offer
📌 Senior Spark / PySpark Data Engineer (Kolkata)
🏢 Web Spiders
📍 Kolkata