14 Aug
|
LH2 AI Labs
|
Bengaluru
14 Aug
LH2 AI Labs
Bengaluru
Job Title: Founding Engineer - Privacy & PII [Data Products]
Location: On-site, Bengaluru
Employment Type: Full-Time About the role
Our entire product rests on one promise: raw enterprise data goes in, and what comes out is safe to sell — nothing identifiable, nothing leaked, provably de-identified, with the audit trail to back the claim. You'll own that promise. You'll build the detection, entity-consistent pseudonymization, vault, and residual-risk gates that convert sensitive enterprise data into a trainable product buyers and regulators can trust.
This is a hands-on founding role and the most trust-critical seat on the team. The difference between a sellable product and a liability is the judgment in this role. Responsibilities
Own the de-identification pipeline end to end: sensitive-entity detection → entity-consistent pseudonymization → residual-risk validation → clean export
Build PII/PHI/secret detection across free text, structured fields, and layout, using and extending tools like Presidio plus custom recognizers
Design and operate the vault: one durable, namespaced surrogate per resolved entity, so the same person maps to the same pseudonym across every system (name, email, employee id, Slack/Jira/Git ids) — without token drift
Build the "rewrite all surfaces" logic that applies surrogates consistently across the dataset
Design and enforce the residual-identifier release gate (regex + NER + secret scanners + sample review) and make explicit where it does and doesn't establish legal anonymization
Enforce the trust boundary: raw stays encrypted inside the customer environment; vault keys never leave; buyers receive only the cleaned product
Build fail-closed behavior (no raw pass-through when the vault or detection is unavailable) and full audit/lineage on every reveal and export
Produce the method cards / provenance records that document de-identification technique and residual-risk posture for buyers
Design detection and compliance profiles as swappable configuration so new verticals (medical/PHI later) extend the system rather than rewrite it
Partner with the Data Products engineer at the handoff point, and with privacy counsel on what the product can honestly claim Must-have skills
6+ years engineering, including direct experience building PII detection, de-identification, tokenization, or privacy-aware data processing (this is the non-negotiable filter — general data-engineering experience does not substitute)
Robust Python; comfort with NLP/NER approaches for entity detection in unstructured text
Deep, demonstrable understanding of the concepts that make or break this product: entity resolution before tokenization, token drift and why consistent surrogates matter, quasi-identifier re-identification risk, and the difference between pseudonymization and anonymization
Experience designing tokenization/pseudonymization systems with a durable mapping store (vault) and format/consistency guarantees
Working knowledge of privacy regimes relevant to enterprise and health data (GDPR/EDPB framing, HIPAA Safe Harbor vs.
Expert
Determination) and how to engineer against them — enough to know what can and can't be claimed
Ability to balance aggressive de-identification against preserving the data's training utility (over-redaction destroys value; under-redaction destroys trust)
Comfortable owning a trust-critical system and its quality bar in a small founding team Nice-to-have skills
Presidio internals, custom recognizers, or comparable NER-based detection at scale
Vault/tokenization backends (HashiCorp Vault Transform, Skyflow, Protecto, or equivalent) and format-preserving encryption
Cloud DLP tooling deployed in-VPC (AWS Comprehend/Macie, Azure PII, Google DLP)
Healthcare de-identification / HIPAA experience (directly relevant to our medical vertical)
Differential privacy or synthetic-data familiarity for high-risk aggregate cases
Data lineage/audit tooling
📌 Founding Engineer - Privacy & PII [ Data Products ] (Bengaluru)
🏢 LH2 AI Labs
📍 Bengaluru