21 Aug
|
Zorba AI
|
Delhi
Data Engineer (Spark/Scala)
About The Role
We are seeking an experienced Data Engineer to design, build, and optimize complex data workflows across on-premises and cloud environments. This role requires deep hands-on expertise in Apache Spark, Databricks, and Scala/PySpark, along with strong SQL and Python skills, to build robust, high-performance data pipelines. You will work extensively on complex on-prem workflows, integrating data across multiple file systems and formats, migrating and modernizing legacy processes, and ensuring productive, reliable data movement across heterogeneous environments.
Key Responsibilities
- Design, develop, and maintain large-scale data pipelines using Apache Spark, Databricks, Scala Spark, and PySpark
- Build and support complex on-premises data workflows, including migration/hybrid on-prem-to-cloud integration patterns
- Integrate data across diverse file systems (on-prem file shares, NAS, HDFS, S3) and formats JSON, Parquet, Fixed-Length, CSV, Excel, Avro
- Write productive, optimized SQL for data extraction, transformation, and loading across relational databases
- Connect to and extract data efficiently from various source databases, tuning queries and pipelines for performance at scale
- Develop and maintain workflow orchestration using Airflow (or similar schedulers) for reliable, monitored pipeline execution
- Write clean, production-grade Python code for data processing, automation, and tooling
- Build and maintain unit/integration tests for data pipelines to ensure data quality and reliability
- Create and maintain clear technical documentation for pipelines, data flows, and system architecture
- Troubleshoot and resolve data pipeline failures, performance bottlenecks, and data quality issues in complex, multi-system workflows
- Collaborate with cross-functional teams (data science, analytics, application engineering) to support downstream data consumption
- Support cloud integration efforts, particularly with Azure, as workloads evolve from on-prem to hybrid/cloud architectures
Required Qualifications
Primary Skills
- Strong hands-on experience with Apache Spark and Databricks for large-scale data processing
- Proficiency with Amazon S3 for data storage and pipeline integration
- Strong SQL skills - query optimization, complex joins, performance tuning
- Proven experience integrating data across various file systems and formats: JSON, Parquet, Fixed-Length, CSV, Excel, Avro, etc.
- Strong knowledge of Scala Spark and PySpark for distributed data processing
- Strong Python programming skills for scripting, automation, and data engineering tasks
- Strong experience connecting to and efficiently extracting data from databases (relational/other),
including performance-conscious extraction strategies
- Demonstrated experience working on complex on-prem data workflows (multi-system integration, legacy system data extraction, hybrid on-prem/cloud pipelines)
- Experience leveraging coding assistant tools and implementing AI agents to enhance development productivity and task execution.
Secondary Skills
- Experience with Azure cloud services (storage, compute, data services)
- Experience with Apache Airflow for workflow orchestration and scheduling
- Experience writing automated tests for data pipelines (unit, integration, data quality checks)
- Strong documentation skills able to clearly document pipelines, data lineage, and technical designs
Good To Have
- Working knowledge of Java
- Familiarity with React for building internal tooling/dashboards
- Experience with Prefect for workflow orchestration
- PBM (Pharmacy Benefit Management) / Healthcare domain knowledge
Data Engineer (Spark/Scala)
Good To Have
- Working knowledge of Java
- Familiarity with React for building internal tooling/dashboards
- Experience with Prefect for workflow orchestration
- PBM (Pharmacy Benefit Management) / Healthcare domain knowledge.
Data Engineer (Spark/Scala)
Good To Have
- Working knowledge of Java
- Familiarity with React for building internal tooling/dashboards
- Experience with Prefect for workflow orchestration
- PBM (Pharmacy Benefit Management) / Healthcare domain knowledge
Skills: scala,pipelines,data,spark
📌 Data Engineer-Spark,Scala (Delhi)
🏢 Zorba AI
📍 Delhi