19 Aug
|
Zorba AI
|
Hyderabad
19 Aug
Zorba AI
Hyderabad
We are looking for an experienced Data Engineer (Spark/Scala) with strong hands-on expertise in Apache Spark, Databricks, Scala, PySpark, Python, and SQL.
The role involves designing and developing large-scale data pipelines across on-premises and cloud environments, working with multiple file systems and data formats, modernizing legacy workflows, and supporting hybrid data architectures.
Key Responsibilities
- Design, develop, and maintain scalable data pipelines using Apache Spark, Databricks, Scala Spark, and PySpark.
- Build and support complex on-premises data workflows and hybrid on-prem-to-cloud data integration solutions.
- Integrate data across HDFS, NAS, on-prem file shares, Amazon S3, and other storage platforms.
- Work with multiple data formats including JSON, Parquet, CSV, Avro, Fixed-Length, and Excel.
- Develop optimized SQL queries for data extraction, transformation, and loading.
- Connect to multiple relational and non-relational databases and implement performance-efficient data extraction strategies.
- Develop and maintain workflow orchestration using Apache Airflow or similar scheduling tools.
- Write clean, production-grade Python code for data processing, automation, and engineering utilities.
- Develop unit, integration, and data-quality tests for data pipelines.
- Troubleshoot pipeline failures, performance bottlenecks, data quality issues, and complex multi-system integration problems.
- Support migration and modernization of legacy on-premises data processes to hybrid/cloud environments.
- Collaborate with Data Scientists, Analysts, Application Engineers, and other stakeholders.
- Create technical documentation covering pipelines, data flows, architecture, and data lineage.
- Support cloud integration initiatives,
particularly across Azure environments.
- Leverage coding assistants and AI agents to improve development productivity and automate engineering tasks.
Required Skills
- Strong hands-on experience with Apache Spark and Databricks.
- Strong experience with Scala/Spark Scala and PySpark.
- Solid Python programming skills.
- Strong SQL, including complex joins, query optimization, and performance tuning.
- Hands-on experience with Amazon S3.
- Experience working with HDFS, NAS, on-prem file systems, and cloud storage.
- Strong experience handling JSON, Parquet, CSV, Avro, Fixed-Length, and Excel data formats.
- Experience extracting data efficiently from multiple databases.
- Strong understanding of complex on-premises data workflows and multi-system integrations.
- Experience building hybrid on-prem/cloud data pipelines.
- Strong troubleshooting and production support skills.
Secondary Skills
- Azure cloud services.
- Apache Airflow or similar workflow orchestration tools.
- Automated unit and integration testing.
- Data quality validation and monitoring.
- Data lineage and technical documentation.
Good to Have
- Working knowledge of Java.
- Experience with Prefect.
- Familiarity with React for internal tools or dashboards.
- Experience using AI coding assistants and AI agents.
- PBM / Pharmacy Benefit Management / Healthcare domain experience.
Preferred Candidate Profile Candidates with strong experience in Scala + Spark + Databricks + PySpark, combined with on-premises data engineering and hybrid cloud integration, will be preferred. Key Skills: Apache Spark, Scala, Spark Scala, Databricks, PySpark, Python, SQL, Amazon S3, HDFS, Airflow, Azure, On-Premise Data Engineering, ETL, Data Pipelines.
Skills: cloud,apache spark,scala
📌 Data Engineer (Spark/Scala) (Hyderabad)
🏢 Zorba AI
📍 Hyderabad