27 Sep
|
Engineering-Enterprise Data Platforms
|
Jaipur
27 Sep
Engineering-Enterprise Data Platforms
Jaipur
We are looking for a Lead Data Architect to design, build, and scale our data
pipelines and entity resolution systems. This role combines deep technical
expertise in data engineering with hands-on experience in AI-assisted tooling,
entity matching, and data integration from diverse sources. You will lead
architectural decisions for our data platform, mentor engineers, and ensure our
pipelines are reliable, scalable, and production-grade.
Key Responsibilities
Architect, build, and maintain robust, scalable data pipelines that
ingest, transform, and serve data from multiple internal and external sources.
Own the end-to-end orchestration of data workflows using tools like
Dagster, ensuring observability, reliability, and maintainability of pipelines.
Design and implement entity resolution workflows — including matching,
merging, and survivorship logic — using tools such as Splink, to produce clean,
deduplicated, golden records.
Build and maintain web scrapers to source data from external providers,
ensuring resilience to source changes, rate limits, and data quality issues.
Integrate and reconcile data coming from multiple, often inconsistent,
sources into unified, trustworthy datasets.
Design and maintain data models and schemas across transactional and
analytical systems, ensuring consistency, scalability, and performance.
Leverage AI/LLM-based tools and techniques to enhance data pipeline
capabilities — e.g., intelligent data extraction, automated data quality
checks, or AI-assisted entity matching.
Define and enforce best practices around pipeline design, testing,
monitoring, and documentation.
Collaborate closely with data engineers, product managers, and other
stakeholders to translate business requirements into scalable data architecture.
Provide technical leadership and mentorship to the data engineering team.
Required Skills & Experience
Solid hands-on experience building and maintaining production-grade
data pipelines at scale.
Practical experience with Dagster (or similar orchestration tools like
Airflow/Prefect) for pipeline orchestration.
Experience with Splink or similar probabilistic/deterministic record
linkage tools for entity matching, merging, and survivorship.
Robust proficiency in Python, including experience writing and
maintaining web scrapers.
Proven experience integrating and maintaining data pipelines that pull
from multiple, heterogeneous data sources.
Experience applying AI/ML tools within data engineering workflows (e.g.,
LLM-assisted data cleaning, extraction, or matching).
Hands-on experience with relational and distributed databases such as
PostgreSQL and Google Cloud Spanner.
Robust understanding of data modeling principles (normalization,
dimensional modeling, schema design) across OLTP and OLAP systems.
Experience with cloud data warehousing platforms such as BigQuery,
Redshift, and cloud platforms (GCP/AWS/Azure).
Strong communication skills and experience working cross-functionally
with engineering and product teams.
Experience with distributed data processing frameworks (e.g., Spark,
Dask).
Familiarity with data governance, lineage, and cataloging tools.
Prior experience in a lead or architect-level role guiding a data
engineering team.
What We're Looking For
A technically strong, hands-on leader who can balance architectural thinking
with the practical grit of debugging a flaky scraper or tuning a matching
algorithm — someone who's comfortable owning both the big picture and the messy
details of real-world data.
📌 Lead Data and AI Architect (Jaipur)
🏢 Engineering-Enterprise Data Platforms
📍 Jaipur