Data Pipeline Engineer-Kafka(SJ-BY) (Kochi)

Data Pipeline Engineer-Kafka(SJ-BY) (Kochi)

09 Aug
|
SE MENTOR SOLUTIONS
|
Kochi

09 Aug

SE MENTOR SOLUTIONS

Kochi

Responsiblities

Build batch-seed and event-tail ingestion per source system, including seed→ tail watermark handoff, idempotent upserts, and dedup ledgers

- Build and operate medallion layers with reprocess-from-Bronze, pipeline orchestration (checkpoints, retry/backoff, DLQ), and full observability

- Build data-quality gates (quarantine / pass-with-flag), quality scoring, and a reconciliation engine covering count, record, and financial reconciliation — financial is zero-tolerance

- Build identity matching combining deterministic rules with probabilistic scoring and confidence bands; deliver deduplication, golden-record materialization, and survivorship rules, calibrating match thresholds with labelled data

- Author and maintain source→ canonical structural mappings and value crosswalks (e.g., collapsing 1,800+ raw employment-status values to ~20 standard ones) as governed, versioned configuration

- Enforce data contracts at the boundary: schema registry, fail-fast validation, and semver-compatible schema evolution What We’re Looking For

- 5+ years building production data pipelines at scale

- Kafka depth: consumers/producers, replay,



DLQ, exactly-once / idempotent processing patterns

- Strong SQL and solid ETL fundamentals

- Java and/or Python in production

- Medallion/lakehouse layering, CDC, watermark/checkpoint patterns, and batch-stream hand-off

- Data-quality frameworks: validation rules, quarantine and re-entry, quality scoring, reconciliation

- Entity resolution / MDM exposure: record matching, dedup, survivorship — via commercial tools (Informatica MDM, Reltio) or custom builds

- Data mapping and crosswalk discipline: profiling messy datasets, authoring governed reference data, config-as-code (YAML/JSON, Git) Bonus Points

- Probabilistic record linkage at depth — blocking/candidate generation, scoring models, threshold calibration (expected at senior level)

- Schema registry experience (Avro/Protobuf)

- Extracting from mainframe or older RDBMS sources with limited CDC support

- Financial reconciliation in finance-adjacent domains

- Perks administration or healthcare domain knowledge

📌 Data Pipeline Engineer-Kafka(SJ-BY) (Kochi)
🏢 SE MENTOR SOLUTIONS
📍 Kochi

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: data pipeline engineer-kafka(sj-by) (kochi) / kochi

Subscribe to this job alert:

Get the latest job offers by email for: data pipeline engineer-kafka(sj-by) (kochi) / kochi