06 Sep
|
Sparix Global
|
India
06 Sep
Sparix Global
India
This is Hybrid role .Job Locations:
Bengaluru, India / Noida, India / Pune, India / Gurgaon, India / Mumbai, India
Skills Required
Primary Skills
Databricks and PySpark; Advanced SQLmedallion/lakehouse; Data vault 2.0
Key Responsibilities
• Own end-to-end technical design of lakehouse pipelines and frameworks: ingestion patterns, medallion layer standards, metadata-driven frameworks, reusable libraries, and naming/coding conventions.
- Lead and grow a squad of data engineers (typically 5–10): task breakdown and estimation, sprint planning with the Scrum Master/EM, code reviews, pairing, and performance feedback.
- Translate business and architecture requirements into technical designs and delivery plans; present and defend design decisions and trade-offs to QBE architects and platform owners.
- Set and enforce engineering quality gates: PR review standards, unit/integration test coverage for pipelines, data quality SLAs, CI/CD promotion criteria, and documentation.
- Own non-functional outcomes — performance, cost optimisation (cluster/DBU governance), security (Unity Catalog permissions, PII handling), reliability, and observability of the platform.
- Manage technical risk: dependency tracking, proof-of-concepts for new patterns (streaming, DLT, Asset Bundles), remediation plans for tech debt, and production incident command for critical issues.
- Coordinate across workstreams — Data Modelers, BDAs, testing, governance — to keep specs, models, code, and test coverage in lockstep; run design authority sessions.
- For the Onshore Lead: front-door for client stakeholders — requirement workshops, steering-committee technical inputs, escalation handling, and onshore–offshore handshake quality.
- For the Offshore Lead: run offshore ceremonies, own sprint delivery and status,
ensure overlap-hours coverage, and manage the offshore–onshore work packaging.
Must-Have Skills & Experience
• 12+ years in data engineering with 5+ years leading teams delivering on Spark platforms; deep, current hands-on Databricks + PySpark (this is a coding lead, not a pure people-manager role).
- Proven architecture-level command of the lakehouse/medallion pattern, Delta Lake, Unity Catalog, and Azure data services (ADLS Gen2, ADF, Key Vault, networking basics for Databricks).
- Track record designing metadata/config-driven ingestion and transformation frameworks used by multiple teams.
- Advanced Spark performance engineering and cost governance at platform scale.
- Strong SDLC leadership: Git branching strategy, CI/CD (Azure DevOps/GitHub Actions), environment strategy, release management.
- Insurance domain experience — P&C; strongly preferred (policy, claims, premium, reinsurance, actuarial data flows).
- Experience running distributed onshore/offshore delivery models with measurable quality outcomes.
Good-to-Have
• DLT, Databricks Asset Bundles, Terraform for Databricks/Azure.
- Streaming architectures (Auto Loader, Structured Streaming, Kafka/Event Hubs).
- Migration experience (on-prem/legacy ETL → Databricks); Informatica/DataStage/SSIS conversion.
- Databricks Professional certification; Azure Solutions Architect (AZ-305) or Data Engineer (DP-203/DP-700).
Qualifications
• Bachelor's/Master's in Computer Science, Engineering, or related field.
Skilled & Communication Skills
• Executive-ready communication — can hold the room with client architects and translate for business stakeholders.
- Decision-making with documented trade-offs; comfortable saying no with alternatives.
- Mentoring culture-builder; raises the squad's bar rather than becoming the bottleneck.
📌 Data Engineering Lead (Databricks / PySpark) (India)
🏢 Sparix Global
📍 India