21 Aug
|
Zensar Technologies
|
Pune
21 Aug
Zensar Technologies
Pune
Description
We are looking for a hands-on Senior Data Engineer with 7 to 9 years of experience to design, build, and optimize large-scale data pipelines and lakehouse solutions on Databricks. The ideal candidate is a strong individual contributor who stays current with the latest Databricks platform capabilities — including Unity Catalog, Lakeflow (Delta Live Tables), Lakebase, and Databricks' expanding AI/agent tooling — and can apply them to solve real-world data engineering problems at scale.
Responsibilities
Key Responsibilities:
- Design, develop, and maintain scalable ETL/ELT pipelines on Databricks using PySpark, Spark SQL, and Delta Lake for batch and streaming workloads.
- Build and manage declarative pipelines using Lakeflow / Delta Live Tables (DLT), including expectations, data quality checks, and change data capture (CDC).
- Implement and manage data governance, access control, lineage, and data sharing using Unity Catalog across multiple workspaces and clouds.
- Optimize Spark jobs and Databricks clusters for performance and cost, leveraging Photon, serverless compute, auto-scaling, and job clustering best practices.
- Design and implement medallion architecture (bronze/silver/gold) data models and lakehouse patterns for analytics and ML consumption.
- Work with Databricks Workflows (Jobs) to orchestrate multi-task pipelines, including dependency management, retries, and monitoring/alerting.
- Integrate Databricks with cloud-native services (AWS/Azure/GCP) such as S3/ADLS/GCS, Kafka/Event Hubs/Kinesis, Glue/ADF, and IAM/Entra ID for secure, automated data flows.
- Apply CI/CD practices for Databricks using Databricks Asset Bundles (DABs), Repos, and Git integration; automate deployments across dev/test/prod.
- Evaluate and adopt newer Databricks capabilities — Lakebase (serverless Postgres on the lakehouse), Unity Catalog Metrics, Genie/Agent Bricks, Mosaic AI,
and real-time/streaming enhancements — and recommend where they add value to existing pipelines.
- Implement data quality, testing, and observability frameworks (e.g., Great Expectations, DLT expectations, Lakehouse Monitoring) to ensure trustworthy, production-grade data.
- Collaborate with data scientists, analysts, and business stakeholders to understand requirements and translate them into robust, reusable data engineering solutions.
- Mentor junior engineers, participate in code reviews, and contribute to engineering best practices, coding standards, and documentation.
- Troubleshoot production data pipeline issues, perform root-cause analysis, and drive continuous improvement in reliability and performance.
Qualifications
Required Skills & Experience
- 7–9 years of overall experience in Data Engineering, with at least 3–4 years of hands-on, production experience on the Databricks platform.
- Solid programming skills in Python and/or Scala, with deep hands-on expertise in PySpark and Spark SQL.
- Solid experience with Delta Lake (ACID transactions, time travel, schema evolution, optimize/vacuum/Z-ordering, liquid clustering).
- Hands-on experience with Lakeflow / Delta Live Tables (DLT) for building declarative, quality-controlled pipelines.
- Working knowledge of Unity Catalog for centralized governance, fine-grained access control, data lineage, and cross-workspace data sharing.
- Experience with Databricks Workflows/Jobs for pipeline orchestration, scheduling, and monitoring.
- Proficiency with at least one major cloud platform (AWS, Azure, or GCP) and its native storage/compute/security services.
- Experience with streaming technologies such as Structured Streaming, Kafka, Event Hubs, or Kinesis.
- Strong SQL skills, including performance tuning, partitioning strategies, and query optimization on large datasets.
- Familiarity with CI/CD for data platforms — Databricks Asset Bundles, Git-based version control, Terraform, and automated testing/deployment pipelines.
- Understanding of data modeling concepts (dimensional modeling, medallion/lakehouse architecture) and data warehousing fundamentals.
- Demonstrated ability to stay current with the Databricks product roadmap (e.g., Unity Catalog enhancements, Lakebase, Genie/Agent Bricks, Mosaic AI, Lakehouse Monitoring, serverless compute) and apply relevant updates to existing systems.
- Strong analytical, debugging, and performance-tuning skills across the Databricks/Spark stack.
- Excellent communication skills with the ability to work directly with cross-functional stakeholders and, where applicable, mentor junior team members.
Good to Have
- Databricks Certified Data Engineer Associate/Professional or Databricks Certified Associate/Professional Developer for Apache Spark certification.
- Exposure to MLOps/MLflow, Mosaic AI, or Databricks' agent/GenAI tooling (Agent Bricks, Genie, AI/BI dashboards).
- Experience with Lakebase or other Postgres/OLTP-on-lakehouse patterns for operational analytics use cases.
- Experience with dbt, Airflow, or similar orchestration/transformation tools alongside Databricks.
- Prior experience in a regulated or high-governance data environment (finance, healthcare, or similar).
- Contributions to internal frameworks, reusable pipeline templates, or engineering best-practice documentation.
📌 DE&A - Sr. Data Engineer - Databricks (Pune)
🏢 Zensar Technologies
📍 Pune