10 Aug
|
ValueMomentum
|
Secunderabad
10 Aug
ValueMomentum
Secunderabad
- Design, develop, and maintain scalable data pipelines using PySpark and related big data technologies.
- Work with large datasets and develop data models for consumption by data scientists and analysts.
- Optimize Spark jobs for better performance and resource management.
- Design and implement data integration workflows between various data sources.
- Troubleshoot and resolve issues related to data pipelines.
- Collaborate with cross-functional teams to understand business requirements and deliver solutions.
- Ensure data quality and cleanliness using validation and transformation techniques.
- Write and maintain productive, scalable code in Python and PySpark.
- Manage data storage, computation, and scaling on cloud platforms like AWS or Azure.
Requirements
- Bachelor's degree in Computer Science, Engineering, or a related field.
- 2 to 4 years of experience with PySpark for data processing on large-scale datasets.
- Solid understanding of Spark architecture,
including RDDs, DataFrames, and Datasets.
- Strong programming experience in Python, including libraries such as pandas, numpy, and matplotlib.
- Experience with Hadoop, Hive, and NoSQL databases (e.g., Cassandra, MongoDB).
- Working knowledge of cloud computing services (e.g., AWS, Azure, or Google Cloud).
- Familiarity with batch and stream processing (using Kafka, Flink, Spark Streaming).
- Strong problem-solving skills and attention to detail.
- Excellent communication and teamwork skills.
Good To Have
- Experience with Apache Airflow or other orchestration tools.
- Familiarity with Docker or Kubernetes for containerized data environments.
- Experience in implementing and managing CI/CD pipelines, focusing on automating, testing, and deploying code.
📌 PySpark Lead (Secunderabad)
🏢 ValueMomentum
📍 Secunderabad