We are looking for an experienced Data Engineer (Spark/Scala) with strong hands-on expertise in Apache Spark, Databricks, Scala, PySpark, Python, and SQL.
The role involves designing and developing large-scale data pipelines across on-premises and cloud environments, working with multiple file systems and data formats, modernizing legacy workflows, and supporting hybrid data architectures.
Key Responsibilities
Design, develop, and maintain scalable data pipelines using Apache Spark, Databricks, Scala Spark, and PySpark.
Build and support complex on-premises data workflows and hybrid on-prem-to-cloud data integration solutions.
Integrate data across HDFS, NAS, on-prem file shares, Amazon S3, and other storage platforms.
Work with multiple data formats including JSON, Parquet, CSV, Avro, Fixed-Length, and Excel.
Develop optimized SQL queries for data extraction, transformation, and loading.
Connect to multiple relational and non-relational databases and implement performance-productive data extraction strategies.
Develop and maintain workflow orchestration using Apache Airflow or similar scheduling tools.
Write clean, production-grade Python code for data processing, automation, and engineering utilities.
Develop unit, integration, and data-quality tests for data pipelines.
Troubleshoot pipeline failures, performance bottlenecks, data quality issues, and complex multi-system integration problems.
Support migration and modernization of legacy on-premises data processes to hybrid/cloud environments.
Collaborate with Data Scientists, Analysts, Application Engineers, and other stakeholders.
Create technical documentation covering pipelines, data flows, architecture, and data lineage.
Support cloud integration initiatives, particularly across Azure environments.
Leverage coding assistants and AI agents to improve development productivity and automate engineering tasks.
Required Skills
Strong hands-on experience with Apache Spark and Databricks.
Strong experience with S
📌 Data Engineer (Pune)
🏢 Zorba AI
📍 Pune