12 Sep
|
CIEL HR
|
Bengaluru
A Python and PySpark Developer is responsible for designing, developing, and optimizing data processing applications and pipelines using Python and the Apache Spark framework
Key Responsibilities:
Develop, test, and maintain scalable data pipelines and ETL processes using Python and PySpark to extract, transform, and load large datasetsOptimize and fine-tune existing PySpark applications and workflows for performance improvements and efficiencyWork with big data technologies such as Hadoop, Hive, or Kafka, and AWS cloud platforms. Experienced in Databricks Lakehouse platformEnsure data quality and integrity throughout the data lifecycle by performing data validation and implementing error-handling mechanismsMonitor and troubleshoot data processing jobs in production environments to ensure reliability and prompt issue resolution
Requirements:
Expertise in handling large-scale datasets and collaborating with data engineers and data scientists to build robust, scalable data solutions for analytics, machine learning, and business intelligenceStrong proficiency in Python and relevant libraries. Expertise in Apache Spark and PySpark, including Spark SQL and DataFramesSolid understanding of data processing concepts, ETL processes, Data warehousing, Quality gates, Data pipeline & reconciliationProficiency in SQL and experience with various database systemsExcellent communication and collaboration skills to work with cross-functional teams. Solid analytical and problem-solving skills for debugging complex data issues.
📌 Python /Pyspark Developer (Bengaluru)
🏢 CIEL HR
📍 Bengaluru