16 Aug
|
LH2 AI Labs
|
Bengaluru
16 Aug
LH2 AI Labs
Bengaluru
Job Title: Founding Engineer - Privacy & PII [Data Products]
Location: On-site, Bengaluru
Employment Type: Full-Time
About LH2 AI Labs
Built by second-time founders who have built and sold companies before, LH2 AI Labs is building the post-training infrastructure for frontier AI models.
We bring private, high-quality institutional datasets and vetted domain experts into frontier AI pipelines across verticals such as coding, computer use, agentic workflows, medical, audio, and more.
For AI to keep progressing, it needs high-quality training data drawn from real production use cases. The public web has already been crawled and trained on—there is limited new signal left there. That is where we come in.
Our vision is to create a world where frontier models can access high-quality data on tap, the same way they access compute today.
About the role
Our entire product rests on one promise: raw enterprise data goes in, and what comes out is safe to sell — nothing identifiable, nothing leaked, provably de-identified, with the audit trail to back the claim. You'll own that promise. You'll build the detection, entity-consistent pseudonymization, vault, and residual-risk gates that convert sensitive enterprise data into a trainable product buyers and regulators can trust.
This is a hands-on founding role and the most trust-critical seat on the team. The difference between a sellable product and a liability is the judgment in this role.
Responsibilities
- Own the de-identification pipeline end to end: sensitive-entity detection → entity-consistent pseudonymization → residual-risk validation → clean export
- Build PII/PHI/secret detection across free text, structured fields, and layout,
using and extending tools like Presidio plus custom recognizers
- Design and operate the vault : one durable, namespaced surrogate per resolved entity, so the same person maps to the same pseudonym across every system (name, email, employee id, Slack/Jira/Git ids) — without token drift
- Build the "rewrite all surfaces" logic that applies surrogates consistently across the dataset
- Design and enforce the residual-identifier release gate (regex + NER + secret scanners + sample review) and make explicit where it does and doesn't establish legal anonymization
- Enforce the trust boundary: raw stays encrypted inside the customer environment; vault keys never leave; buyers receive only the cleaned product
- Build fail-closed behavior (no raw pass-through when the vault or detection is unavailable) and full audit/lineage on every reveal and export
- Produce the method cards / provenance records that document de-identification technique and residual-risk posture for buyers
- Design detection and compliance profiles as swappable configuration so new verticals (medical/PHI later) extend the system rather than rewrite it
- Partner with the Data Products engineer at the handoff point, and with privacy counsel on what the product can honestly claim
Must-have skills
- 6+ years engineering, including direct experience building PII detection, de-identification,
tokenization, or privacy-aware data processing (this is the non-negotiable filter — general data-engineering experience does not substitute)
- Solid Python; comfort with NLP/NER approaches for entity detection in unstructured text
- Deep, demonstrable understanding of the concepts that make or break this product: entity resolution before tokenization, token drift and why consistent surrogates matter, quasi-identifier re-identification risk, and the difference between pseudonymization and anonymization
- Experience designing tokenization/pseudonymization systems with a durable mapping store (vault) and format/consistency guarantees
- Working knowledge of privacy regimes relevant to enterprise and health data (GDPR/EDPB framing, HIPAA Safe Harbor vs. Expert Determination) and how to engineer against them — enough to know what can and can't be claimed
- Ability to balance aggressive de-identification against preserving the data's training utility (over-redaction destroys value; under-redaction destroys trust)
- Comfortable owning a trust-critical system and its quality bar in a small founding team
Nice-to-have skills
- Presidio internals, custom recognizers, or comparable NER-based detection at scale
- Vault/tokenization backends (HashiCorp Vault Transform, Skyflow, Protecto, or equivalent) and format-preserving encryption
- Cloud DLP tooling deployed in-VPC (AWS Comprehend/Macie, Azure PII, Google DLP)
- Healthcare de-identification / HIPAA experience (directly relevant to our medical vertical)
- Differential privacy or synthetic-data familiarity for high-risk aggregate cases
- Data lineage/audit tooling
📌 Founding Engineer - Privacy & PII [ Data Products ] (Bengaluru)
🏢 LH2 AI Labs
📍 Bengaluru