Data Pipeline & Ingestion Engineer (Senior / Mid)
Java Data Pipeline – Kafka
Experience: 5-12 years | Job Mode: Work From Office | www.bytestrone.com |
[email protected]
Location : Bytestrone India Pvt. Ltd.
201, Lulu IT Cyber Twin Tower 1,Smart City SEZ, Kakkanad -Kochi -Kerala
*Senior/Lead and Mid-level openings · ODL Program*
About the Role
You will build and operate the data backbone of ODL: bulk and streaming ingestion from legacy source systems, medallion-layered storage (Bronze/Silver/Gold), identity resolution and golden-record consolidation, source-to-canonical mapping and crosswalks, and the data-quality and reconciliation gates that prove data is complete and correct before it is published. This is the volume engine of the program — every current client onboarded flows through the pipelines you build.
What You’ll Do
- Build batch-seed and event-tail ingestion per source system, including seed→ tail watermark hand-off, idempotent upserts, and dedup ledgers
- Build and operate medallion layers with reprocess-from-Bronze,
pipeline orchestration (checkpoints, retry/backoff, DLQ), and full observability
- Build data-quality gates (quarantine / pass-with-flag), quality scoring, and a reconciliation engine covering count, record, and financial reconciliation — financial is zero-tolerance
- Build identity matching combining deterministic rules with probabilistic scoring and confidence bands; deliver deduplication, golden-record materialization, and survivorship rules, calibrating match thresholds with labelled data
- Author and maintain source→ canonical structural mappings and value crosswalks (e.g., collapsing 1,800+ raw employment-status values to ~20 standard ones) as governed, versioned configuration
- Enforce data contracts at the boundary: schema registry, fail-fast validation, and semver-compatible schema evolution
What We’re Looking For
- 5+ years building production data pipelines at scale
- Kafka depth
: consumers/producers, replay, DLQ, e
📌 Java Data Pipeline – Kafka (Kochi)
🏢 Bytestrone
📍 Kochi