24 Sep
|
ValueMomentum
|
Coimbatore
24 Sep
ValueMomentum
Coimbatore
Job Summary
We are seeking a skilled and motivated PySpark Engineer to design, develop, and optimize large-scale data processing solutions using PySpark and distributed computing technologies.
The ideal candidate will have strong hands-on experience in PySpark, Python, SQL, and ETL, along with strong analytical, debugging, and problem-solving skills.
The candidate should also have a strong interest in learning and adapting to new technologies as required by business and project needs.
Pre-Screening Questionnaire Mandatory
Candidates are required to complete the PySpark Engineer Pre-Screening Questionnaire as part of the application process.
Questionnaire Link:
https://forms.cloud.microsoft/r/JuFHXi1Dgz
Please complete the questionnaire after applying for this position.
Roles & Responsibilities
- Design, develop, and maintain scalable data processing applications using PySpark.
- Build and optimize ETL/ELT pipelines for structured and unstructured data.
- Process large datasets efficiently using Apache Spark.
- Integrate data from relational databases, cloud storage, APIs, and streaming platforms.
- Optimize Spark jobs for performance, scalability, and cost efficiency.
- Collaborate with Data Engineers, Data Scientists, Business Analysts, and Architects to deliver business solutions.
- Develop reusable frameworks and follow coding standards for data engineering projects.
- Perform data validation and data quality checks.
- Troubleshoot production issues and resolve performance bottlenecks.
- Monitor production jobs and ensure smooth execution.
- Participate in code reviews and maintain technical documentation.
- Work closely with cross-functional teams throughout the software development lifecycle.
Required Skills
- 47 years of experience in Data Engineering.
- Minimum 3 years of hands-on experience in PySpark development.
- Strong Python programming skills.
- Hands-on experience with Apache Spark, Spark SQL, Spark DataFrames, and RDDs.
- Strong SQL and ETL development experience.
- Positive understanding of distributed computing concepts.
- Experience with AWS and/or Azure cloud platforms.
- Familiarity with Git version control.
- Experience with workflow orchestration tools such as Airflow or Azure Data Factory.
- Good understanding of Data Warehousing concepts and Dimensional Modeling.
- Excellent analytical, debugging, and problem-solving skills.
- Willingness to learn and adapt to recent technologies as required is mandatory.
Preferred Skills
- Databricks
- Delta Lake
- Lakehouse Architecture
- Docker
- Kubernetes
- CI/CD Pipelines
- DevOps practices for Data Engineering
Qualification Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, or a related discipline.
Location
Hyderabad / Pune / Coimbatore / Bengaluru
Notice Period
Immediate to 15 days preferred
📌 Pyspark Developer (Coimbatore)
🏢 ValueMomentum
📍 Coimbatore