30 Sep
|
Blumetra Solutions
|
Hyderabad
30 Sep
Blumetra Solutions
Hyderabad
Job Title: Databricks – Lead Data Engineer
Work Location: Hyderabad (Hybrid)
Experience: 7 – 10 years
Role Summary
We are looking for a hands-on Databricks Lead to design, build and take to production governed financial data products on an Azure Databricks lakehouse. You will own the Databricks solution end to end – file-based ingestion, medallion pipelines, hierarchy modelling, row-level security, zero-outage publishing and performance – and lead a small engineering team while working closely with source-system owners, finance stakeholders, and security and platform teams.
What You Will Build A typical engagement publishes multi-dimensional financial data (for example a management P&L; with entity, department, project, product, version, period and account dimensions) from an enterprise planning system as a single, row-secured consumption table in Databricks, refreshed several times a day. The quality bar:
- Exact reconciliation to source control totals; no data loss.
- Full fidelity of hierarchies, alternate hierarchies, aliases and attributes.
- Zero consumer outage during refresh; a failed run never replaces the last positive data.
- Row-level security that matches the source system's access rules.
- Sub-second response for agreed benchmark queries.
- Additive schema changes absorbed without code changes; full rebuild from Git.
Key Responsibilities
Architecture and technical leadership
- Own the Databricks solution design and drive open decisions with stakeholders: landing zone, publish mechanism, orchestration pattern, compute, retention and recovery targets.
- Translate integration specifications into buildable designs, data contracts and a delivery plan; lead design reviews with integration architects and platform teams.
- Lead, mentor and review the work of 2–4 data engineers; set coding, testing and documentation standards.
- Raise risks early (data volume, security parity, BI connectivity) and propose options with trade-offs.
Platform setup and governance (Unity Catalog)
- Set up catalogs, schemas, storage credentials, external locations and volumes for landing zones and medallion layers across dev, test and prod.
- Define the grant model with account-level groups synced from Microsoft Entra ID; configure service principals for jobs and CI/CD.
- Work with Azure network teams on private endpoints for ADLS and Network Connectivity Configurations for serverless compute.
Ingestion and orchestration
- Build manifest-gated, file-arrival-triggered Databricks Workflows with task dependencies, task values, condition-task publish gates, queueing, retries, timeouts and notifications.
- Implement Bronze ingestion from delimited files with checksum and row-count verification against a JSON manifest.
- Guarantee idempotent reruns and safe handling of overlapping or late runs.
Data modelling and transformation (Silver / Gold)
- Build PySpark logic to validate and flatten ragged parent-child hierarchies, detecting cycles, orphans, multi-parent members and non-unit weights.
- Handle alternate hierarchies through configuration, as flattened tables or bridge tables.
- Pivot long-form attributes and aliases into columns; build a contract-driven Gold generator that absorbs additive schema changes automatically.
- Design wide, denormalized consumption tables, including leaf/consolidated grain flags to prevent double counting.
Data quality and reconciliation
- Implement blocking quality checks (contract, referential integrity, hierarchy, typing, enrichment, security) and exact reconciliation to source control totals.
- Build control tables (run log, check results, reconciliation, schema contract) and an operations dashboard.
Security
- Implement source-system security parity: map source users to Entra identities, resolve group and member-level rights into a security lookup table, and enforce them with a Unity Catalog row-filter function.
- Keep data and security consistent during publish; define break-glass access, ownership and audit monitoring.
- Advise BI teams on per-user connectivity (SSO / DirectQuery) so row-level security holds end to end.
Publishing and performance
- Implement zero-outage publishing with atomic Delta overwrites or blue/green swaps, plus rollback via time travel.
- Tune for sub-second queries: liquid clustering, data-skipping statistics, predictive optimization, Photon,
SQL Warehouse sizing and row-filter cost; run benchmark tests with real access profiles.
DevOps, testing and operations
- Package everything as Databricks Asset Bundles with CI/CD (GitHub Actions or Azure DevOps) across dev, test and prod.
- Own the test strategy: unit tests, contract and fixture tests, integration, SIT reconciliation, security UAT and performance tests.
- Write the runbook (rerun, rollback, full rebuild, schema change) and hand over to production support.
Collaboration
- Agree file contracts, manifests and control totals with source-system administrators.
- Work with finance users to define benchmark queries and validate numbers during UAT and parallel runs.
- Report progress, risks and decisions to stakeholders and delivery leadership.
Required Skills
Area
Databricks core
Unity Catalog
Fine-grained security
Delta Lake
Performance
PySpark and SQL
Python engineering
Azure
DevOps
Data modelling
Data quality
Must Have Delivered At Least One Of
- A production Databricks pipeline with Unity Catalog row-level security in a regulated or finance setting.
- A file-based, event-triggered ingestion framework with manifest or control-file validation.
- A financial planning or EPM data integration into a lakehouse or warehouse.
Preferred Skills, Qualifications And Attributes
Nice to have
- Enterprise planning / EPM platforms: multi-dimensional cubes, dimensions and hierarchies, rules and consolidations, member-level security, and exporting data from them.
- FP&A; domain knowledge: management P&L;, budget and forecast versions, adjustments and allocations, spend by product or project, GAAP vs non-GAAP.
- Power BI with Databricks: DirectQuery, Entra SSO pass-through, impact of import models on row-level security.
- Lakeflow Declarative Pipelines (DLT) and expectations, Auto Loader, file events on external locations.
- Databricks system tables for cost and job monitoring; Lakeview dashboards and SQL alerts.
- Experience in regulated environments (SOX controls, audit trails, change management).
Certifications (any is a plus)
- Databricks Certified Data Engineer Professional
- Databricks Certified Data Engineer Associate
- Microsoft Certified: Azure Data Engineer Associate or Azure Solutions Architect Expert
Skills: security,databricks,azure,medallion architecture,data,pyspark,unity
📌 Databricks Lead (Hyderabad)
🏢 Blumetra Solutions
📍 Hyderabad