06 Aug
|
Zorba AI
|
Gurugram
Role Summary
We are seeking an experienced Data Engineer with strong expertise in Databricks and up-to-date Lakehouse architecture to design, develop, and optimize scalable data pipelines. The ideal candidate will have hands-on experience with Apache Spark (PySpark), Delta Lake, Python, SQL, and GCP, along with a strong understanding of data engineering best practices, performance optimization, and enterprise-scale data platforms.
Key Responsibilities
- Design, develop, and maintain scalable batch and streaming data pipelines using Databricks.
- Build robust ETL/ELT data pipelines using PySpark and Spark SQL.
- Develop and manage Delta Lake tables, ensuring ACID compliance, schema evolution, and efficient data management.
- Implement and maintain Lakehouse Medallion Architecture (Bronze, Silver, Gold).
- Perform data ingestion from multiple sources, including APIs, Kafka, cloud storage, and relational databases.
- Design and develop data models to support analytics, reporting, and business intelligence.
- Ensure data quality through validation, reconciliation, and monitoring processes.
- Optimize Spark workloads using partitioning, caching, query tuning, and cluster optimization techniques.
- Implement data governance and security using Unity Catalog or equivalent governance tools.
- Integrate data pipelines with orchestration tools such as Apache Airflow or Cloud Composer.
- Collaborate with business stakeholders, architects, and cross-functional teams to understand requirements and deliver scalable solutions.
- Troubleshoot production issues and continuously improve pipeline performance and reliability.
Mandatory Skills Databricks & Big Data
- Strong hands-on experience with Databricks (GCP preferred).
- Expertise in Apache Spark (PySpark).
- Experience with Delta Lake, including MERGE, UPSERT, Time Travel, and schema evolution.
- Strong understanding of Lakehouse Architecture.
Programming
- Strong proficiency in Python.
- Advanced SQL skills.
Data Engineering
- ETL/ELT pipeline development.
- Batch and streaming data processing.
- Data modeling (Dimensional Modeling/Lakehouse).
Cloud Platform
- Hands-on experience with Google Cloud Platform (GCP).
- Experience with services such as:
- BigQuery
- Cloud Storage
- Cloud Composer
- Dataflow (preferred)
Good To Have Skills
- Apache Airflow / Cloud Composer.
- Kafka or other real-time streaming platforms.
- CI/CD using Git, GitHub Actions, or Azure DevOps.
- Infrastructure as Code (Terraform).
- Docker and Kubernetes.
- Exposure to BI tools such as Power BI or Tableau.
- Experience working in BFSI or other enterprise data environments.
Skills: databricks,gcp,sql,apache spark,delta lake
📌 Data Engineer – Databricks (Gurugram)
🏢 Zorba AI
📍 Gurugram