11 Sep
|
Zinnov Management Consulting
|
Bengaluru
11 Sep
Zinnov Management Consulting
Bengaluru
Zinnov is hiring for the role of Professional, Sr.
Databricks Platform
Engineer on behalf of our Global MNC MedTech client company - a company where technology sits at the centre of everything: powering how the core business operates day to day and shaping the products and digital solutions that will define the future of patient care. This role is part of the company's establishment of its new India Global Capability Centre (GCC) in Bengaluru, which will drive enterprise technology, digital transformation and innovation for global operations.
The role works closely with data architecture, engineering, analytics, security, governance, and cloud infrastructure teams to operate the Databricks Lakehouse as a secure, scalable, observable, and cost-productive enterprise platform.
Role & responsibilities
- Design and maintain scalable batch and streaming pipelines using PySpark, Spark SQL, Delta Lake, Auto Loader, and Databricks Lakeflow Declarative Pipelines (formerly Delta Live Tables) or equivalent.
- Architect bronze/silver/gold Lakehouse layers with schema evolution, incremental processing, change data capture, data curation, and reusable ingestion patterns.
- Administer Databricks workspaces, compute, cluster policies, SQL Warehouses, serverless capabilities, jobs/workflows, identities, and environment configuration.
- Implement Unity Catalog for catalogs, schemas, lineage, data discovery, fine-grained permissions, row/column controls, and governed data sharing.
- Optimize Spark workloads and platform cost through partitioning, liquid clustering or Z-ordering, caching, Photon, autoscaling, adaptive query execution, and workload right-sizing.
- Automate platform and workload deployments using Git, Databricks Asset Bundles, Terraform, Azure DevOps/GitHub Actions/Jenkins, and CI/CD practices.
- Integrate Databricks with Azure/AWS services, SAP and ERP sources, APIs, streaming platforms, BI tools, and downstream data science consumers.
- Implement secrets management, encryption, private networking, identity federation/service principals, role-based access, audit logging, and enterprise security controls.
- Establish monitoring, alerting, data-quality controls, and production support for platform health and pipeline reliability, leading root-cause analysis of complex failures.
- Enable analytics and ML teams with curated datasets, feature pipelines, MLflow-supported workflows, model-serving integration, and documented engineering standards.
Preferred candidate profile 6+ years of progressive data engineering or platform engineering experience with substantial hands-on experience operating production Databricks workloads.
Advanced Apache Spark skills using PySpark and Spark SQL, including distributed processing concepts and performance tuning.
Deep hands-on experience with Delta Lake and Lakehouse design, including ACID transactions, schema enforcement/evolution, time travel, optimization,
and incremental processing.
Experience with Databricks administration including workspaces, compute/cluster policies, Jobs/Workflows, SQL Warehouses, and Unity Catalog.
Strong Python and SQL development skills with experience building modular, testable, production-grade data pipelines.
Hands-on Azure or AWS experience including object storage (ADLS/S3), identity, networking, security, monitoring, and service integration.
Experience implementing Git-based CI/CD and infrastructure-as-code using Terraform and Databricks Asset Bundles or equivalent.
Demonstrated ability to troubleshoot complex platform, Spark, data, and production reliability issues.
Preferred
Experience with Structured Streaming, Kafka, Event Hubs, Kinesis, or other real-time ingestion patterns.
Experience integrating SAP S/4HANA, BW, or other enterprise ERP data into a Lakehouse architecture.
Familiarity with MLflow, feature engineering/feature stores, Mosaic AI, and supporting ML/AI workloads on Databricks.
Exposure to data cataloging and governance platforms such as Collibra, Microsoft Purview, or Alation.
MedTech, Life Sciences, healthcare, or other regulated-industry experience with familiarity in GxP and ALCOA+ data-integrity expectations.
Databricks Certified Data Engineer Professional or equivalent certification preferred; cloud data engineering certification is an advantage.
Additional Information
Language: English proficiency required; additional regional languages are a plus.
Travel: Limited domestic or international travel may be required, typically less than 10%.
📌 Professional, Sr. Databricks Platform Engineer (Bengaluru)
🏢 Zinnov Management Consulting
📍 Bengaluru