14 Aug
|
LH2 AI Labs
|
Bengaluru
14 Aug
LH2 AI Labs
Bengaluru
Job Title: Founding Engineer - Data Products
Location: On-site, Bengaluru
Employment Type: Full-Time About the role
We're building the platform that turns raw enterprise data — Slack, Jira, GitHub, email, docs, tickets — into structured, machine-trainable products for AI labs. You'll own the data foundation end to end: how we ingest from many enterprise sources, normalize the mess into one common shape, resolve identities across systems, and assemble the structured "decision" units we sell.
This is a hands-on founding role on a small team. You'll make the core architectural calls, set the engineering patterns the rest of the team builds on, and get a real product shipping inside customer cloud environments fast. Responsibilities
Own the ingestion → normalization → entity-resolution → product-assembly pipeline architecture
Build and operate production connectors across multiple SaaS/API sources (starting Slack, Jira, GitHub; expanding to Google Workspace, Teams, Notion, Confluence, CRM)
Implement reliable incremental sync, pagination, retries, backfills, and checkpointing
Design the medallion data model (raw → normalized entities → sellable products) and the transformation layers between them
Build tiered entity resolution — authoritative and deterministic cross-system identity joins first, probabilistic matching later — with confidence and evidence retained
Assemble structured, citation-backed "task/decision" units from resolved data, seeded from concrete outcomes (merged PRs, closed tickets)
Handle schema evolution, deduplication, late/deleted data, and per-tenant workflow differences via config, not code forks
Build pipelines that are idempotent,
replayable, observable, and reproducible (versioned output manifests)
Establish data-engineering standards and patterns for the team
Work closely with the Privacy/PII founding engineer so the pipeline hands off cleanly at the trust boundary Must-have skills
6+ years of data/software engineering, including owning a production pipeline that ingested from multiple external APIs with incremental sync, retries, and schema drift (this is the non-negotiable – not "I've used a pipeline," but "I built and operated one")
Strong Python and SQL
Solid ETL/ELT experience across multiple SaaS platforms and APIs
Deep understanding of OAuth, pagination, rate limits, incremental/delta sync, backfills, idempotency, and failure recovery
Experience designing data models, normalization layers, and entity resolution across heterogeneous sources
Experience with a modern warehouse/lakehouse (Snowflake, BigQuery, Databricks, Redshift, or equivalent) and orchestration (Airflow, Dagster, Prefect, or equivalent)
Comfortable owning architectural decisions and setting patterns in a small founding team
Experience operating data workloads in cloud environments, ideally customer VPC / private-cloud deployments
Nice-to-have skills
Connector platforms (Airbyte, Meltano, Fivetran) and building custom API connectors
Entity-resolution tooling (Splink, Zingg) and record-linkage fundamentals
Identity/directory/HR data sources as an ER seed
Data lineage/catalog tooling (OpenLineage or similar)
Handling large volumes of semi-structured/unstructured data
Familiarity with AI training-data workflows (SFT, RLHF, evals) and what labs actually buy
📌 Founding Engineer - Data Products (Bengaluru)
🏢 LH2 AI Labs
📍 Bengaluru