01 Sep
|
Finarb
|
Kolkata
Job Description
Location : [Remote / Hybrid — Kolkata, Hyderabad, Bangalore]
n
Type : Full-time
n
Level : Senior
n
n
About the role
n
n
We're looking for a Senior Data Engineer to design, build, and operate the data platform that powers our analytics and reporting. You'll own data pipelines and models end to end — from source ingestion through a Lakehouse medallion architecture to the semantic layer that business users rely on. This is a long-term build-and-run role: you'll ship new data products, keep existing ones healthy, and continuously raise the quality and performance bar of the platform.
n
n
We run on Microsoft Fabric (Lakehouse, Delta, DirectLake Power BI), so Fabric experience is a strong plus — but we care more about deep, transferable data engineering fundamentals than any single vendor stack.
n
n
What you'll do
n
n
n
- Design and build batch and incremental data pipelines in PySpark across Bronze/Silver/Gold layers.
n
n
n
n
- Model data for analytics — dimensional / star schemas, slowly changing dimensions, conformed dimensions, and well-partitioned Delta tables.
n
n
n
n
- Own the semantic layer: build and maintain Power BI models and DAX measures (DirectLake), and partner with analysts on what the business needs.
n
n
n
n
- Build data quality and validation into everything — row/column parity, schema checks, freshness, null/range/dedup rules — so issues are caught before stakeholders see them.
n
n
n
n
- Tune performance and cost — partitioning, file compaction, broadcast joins,
and eliminating common Spark anti-patterns.
n
n
n
n
- Operate what you build: monitoring, logging, alerting, incident response, and clean CI/CD via Azure DevOps.
n
n
n
n
- Collaborate in a shared codebase with strong Git discipline — branches, PRs, code review — and mentor more junior engineers.
n
n
n
n
- Translate ambiguous business asks into reliable, documented, maintainable data products.
n
n
n
What we're looking for
n
n
Required
n
n
n
- 3+ years in data engineering, with significant Spark / PySpark at production scale.
n
- Strong SQL and dimensional data modeling (Kimball-style star schemas, SCD patterns, surrogate keys).
n
- Solid grasp of lakehouse / medallion architecture — what belongs in each layer and why — and Delta Lake (MERGE/upserts, time travel, OPTIMIZE, partitioning).
n
- Comfort building incremental, idempotent pipelines rather than full-reload jobs.
n
- Python engineering fundamentals — testing, modularity, reusable libraries.
n
- Git-based collaboration and CI/CD (Azure DevOps, GitHub Actions, or similar).
n
- A quality- and ownership-first mindset: you instrument, validate, and monitor your own work.
n
- n
- Git-based collaboration and CI/CD (Azure DevOps, GitHub Actions, or similar).
n
- A quality- and ownership-first mindset: you instrument, validate, and monitor your own work.
n
n
n
Preferred
n
n
n
- Microsoft Certified: Fabric Data Engineer Associate (DP-700) — strongly preferred.
n
- Hands-on Microsoft Fabric (Lakehouse, OneLake, Notebooks, Data Pipelines) and Power BI / DirectLake + DAX.
n
- Agentic / AI-assisted development — productive with LLM coding agents and tooling (e.g. Claude Code, MCP servers, custom agents/skills) to accelerate engineering work while keeping a human-in-the-loop quality bar.
n
- Python data-quality and testing tooling — pytest, chispa, Excellent Expectations.
n
- Spark performance tuning at scale.
n
- Streaming or near-real-time ingestion experience.
n
- Exposure to healthcare / pharmacy / pricing data domains.
n
n
n
What success looks like
n
n
n
- You ship reliable data products that pass review on the first or second pass and need little rework.
n
- The pipelines you own are observable, well-tested, and rarely page anyone at night.
n
- You leave the platform cleaner than you found it — better patterns, fewer anti-patterns, faster onboarding for the next engineer.
n
n
n
Tech you'll work with
n
n
Microsoft Fabric (Lakehouse, OneLake, Notebooks, Data Pipelines) · PySpark · Delta Lake · Power BI / DirectLake · DAX · SQL · Python (pytest, chispa, Great Expectations) · Azure DevOps
n
n
n
n
- n
📌 Senior Data Engineer (Kolkata)
🏢 Finarb
📍 Kolkata