05 Sep
|
ValueMomentum
|
Secunderabad
05 Sep
ValueMomentum
Secunderabad
Design, develop, and maintain scalable data pipelines using PySpark and related big data technologies.
Work with large datasets and develop data models for consumption by data scientists and analysts.
Optimize Spark jobs for better performance and resource management.
Design and implement data integration workflows between various data sources.
Troubleshoot and resolve issues related to data pipelines.
Collaborate with cross-functional teams to understand business requirements and deliver solutions.
Ensure data quality and cleanliness using validation and transformation techniques.
Write and maintain effective, scalable code in Python and PySpark.
Manage data storage, computation, and scaling on cloud platforms like AWS or Azure.
Requirements
Bachelor's degree in Computer Science, Engineering, or a related field.
2 to 4 years of experience with PySpark for data processing on large-scale datasets.
Solid understanding of Spark architecture,
including RDDs, DataFrames, and Datasets.
Robust programming experience in Python, including libraries such as pandas, numpy, and matplotlib.
Experience with Hadoop, Hive, and NoSQL databases (e.g., Cassandra, MongoDB).
Working knowledge of cloud computing services (e.g., AWS, Azure, or Google Cloud).
Familiarity with batch and stream processing (using Kafka, Flink, Spark Streaming).
Robust problem-solving skills and attention to detail.
Excellent communication and teamwork skills.
Good To Have
Experience with Apache Airflow or other orchestration tools.
Familiarity with Docker or Kubernetes for containerized data environments.
Experience in implementing and managing CI/CD pipelines, focusing on automating, testing, and deploying code.
📌 Pyspark Lead Secunderabad
🏢 ValueMomentum
📍 Secunderabad