06 Sep
|
Sparix Global
|
India
06 Sep
Sparix Global
India
Skills Required
Primary Skills
Databricks and PySpark; Advanced SQLmedallion/lakehouse; Data vault 2.0
Work Location: Remote (India)
- Candidates should be flexible to work from an EXL office whenever business requires. While the role is primarily remote, it is not a permanent work-from-home position.
Key Responsibilities
• Design, build, and productionise batch and incremental ETL/ELT pipelines using Databricks (notebooks, Jobs/Workflows, DLT), PySpark, and Spark SQL across bronze/silver/gold Delta Lake layers.
- Implement the medallion architecture with Delta Lake features — ACID merges/upserts (SCD1/SCD2), schema evolution/enforcement, time travel, OPTIMIZE/Z-ORDER, vacuuming — and manage data with Unity Catalog.
- Orchestrate workloads via Databricks Workflows and/or Azure Data Factory; parameterise pipelines, implement restartability, checkpointing, and idempotent re-runs.
- Tune Spark for performance and cost: partitioning strategy, join/broadcast optimisation, skew handling, caching, cluster sizing/autoscaling, Photon usage, and job-level cost monitoring.
- Build data quality checks (expectations, reconciliation against source, row/measure-level controls) and implement failure alerting, logging, and lineage-friendly design.
- Ingest from diverse sources — RDBMS (SQL Server/Oracle), files (CSV/Parquet/JSON/XML), APIs, and streaming/CDC feeds — into ADLS Gen2 landing zones.
- Contribute to CI/CD for data: Git-based development, pull-request reviews, Azure DevOps/GitHub Actions pipelines, environment promotion, and infrastructure/config as code.
- Collaborate with Data Modelers, BDAs, testers, and onshore leads in Agile ceremonies; convert mapping specifications into robust, reviewed code; mentor junior engineers.
- Provide L3 production support for owned pipelines:
triage incidents, perform root-cause analysis, and drive permanent fixes within SLAs.
Must-Have Skills & Experience
• 9–12 years overall; robust recent hands-on Databricks and PySpark delivery experience (3+ years Databricks preferred).
- Expert-level Spark (DataFrame API, Spark SQL, window functions, UDF trade-offs) and advanced Python for data engineering.
- Advanced SQL — complex transformations, performance tuning, analytical/window queries on large volumes.
- Delta Lake internals and medallion/lakehouse architecture in production.
- Azure data stack: ADLS Gen2, Azure Data Factory, Key Vault, and Databricks administration basics (clusters, pools, jobs, secrets).
- Data warehousing concepts: star schemas, fact/dimension loading patterns, SCD handling, reconciliation.
- Git-based SDLC with CI/CD exposure (Azure DevOps or GitHub).
- Insurance or financial-services data experience — ideally P&C; (policy/claims/premium/reinsurance structures).
Good-to-Have
• Delta Live Tables, Unity Catalog governance, Databricks Asset Bundles.
- Streaming (Structured Streaming, Auto Loader, Event Hubs/Kafka) and CDC tools.
- dbt on Databricks; Airflow; Terraform.
- Guidewire, Duck Creek, or London-market source systems; Snowflake or Synapse coexistence.
- Databricks Data Engineer Associate/Professional certification.
Qualifications
• Bachelor's degree in Computer Science, Engineering, or equivalent practical experience.
Professional & Communication Skills
• Clear written and spoken English; comfortable presenting design decisions to onshore architects and client stakeholders.
- Proven offshore-delivery discipline: crisp status reporting, proactive risk flagging, dependable overlap-hours availability.
- Ownership mindset — drives issues to closure without follow-up.
📌 Senior Data Engineer (Databricks / PySpark) (India)
🏢 Sparix Global
📍 India