04 Sep
|
Zorba AI
|
Bengaluru
04 Sep
Zorba AI
Bengaluru
We are looking for an experienced Data Engineer with 6+ years of experience in designing, developing, deploying, monitoring, and optimizing scalable batch and streaming data pipelines. The ideal candidate should have solid hands-on expertise in Scala, Apache Spark, Kafka, Airflow, BigQuery, Dataproc, and GCP.
Key Responsibilities
- Design, develop, deploy, monitor, and optimize ETL/ELT workflows for both batch and real-time/streaming data processing.
- Develop scalable data processing solutions using Scala and Apache Spark.
- Build and manage streaming pipelines using Kafka, including handling high-volume data and back-pressure scenarios.
- Develop and maintain Airflow DAGs for complex data orchestration, dependency management, scheduling, retries, and idempotent execution.
- Work extensively with Google Cloud Platform (GCP) services including BigQuery and Dataproc.
- Implement Medallion Architecture (Bronze, Silver, Gold) with appropriate data quality validations, transformations, aggregations, and summarization.
- Perform performance tuning and optimization of Spark jobs, Kafka pipelines, BigQuery queries, and ETL workflows.
- Implement monitoring and observability mechanisms to identify and troubleshoot pipeline failures, performance issues, and data-quality problems.
- Create and maintain runbooks and operational documentation for production support and continuous optimization.
- Troubleshoot runtime data issues independently and ensure timely resolution of production incidents.
- Follow best practices for data governance, security, access control, and information security.
Required Skills
- 6+ years of experience as a Data Engineer.
- Strong hands-on experience with Scala and Apache Spark.
- Strong knowledge of Kafka and streaming data processing.
- Experience with Apache Airflow and workflow orchestration.
- Hands-on experience with Google Cloud Platform (GCP).
- Strong experience with BigQuery and Dataproc.
- Experience building both batch and streaming ETL pipelines.
- Good understanding of Medallion Architecture and data quality frameworks.
- Strong understanding of Spark performance optimization and distributed data processing.
- Knowledge of streaming concepts such as back-pressure handling, fault tolerance, checkpointing, and recovery.
- Experience with monitoring, observability, troubleshooting, and production support.
- Good understanding of idempotency, dependency handling, retries, and failure recovery in Airflow.
Good to Have
- Exposure to data governance and information security.
- Experience with GCP data services beyond BigQuery and Dataproc.
- Experience implementing automated data quality checks and reconciliation.
- Knowledge of CI/CD and infrastructure-as-code practices.
- Experience creating production runbooks and operational standards.
Ideal Candidate A strong candidate should be capable of owning end-to-end data pipelines, from development and deployment through production monitoring, troubleshooting, performance optimization, and continuous improvement, across both batch and real-time data environments. Skills: kafka,gcp,scala,spark
📌 Data Engineer – Scala | Kafka | Spark | GCP (Bengaluru)
🏢 Zorba AI
📍 Bengaluru