The hardest problem in our system isn't retrieval. It's deciding when two mentions are the same person.
A full name in one document. An abbreviation in another. A transliteration in a third script. An honorific. A description with no name at all. Everything above this depends on getting it right. Getting it wrong produces something worse than an error: a confident, well-sourced picture of a person who doesn't exist.
We're hiring one person to own it.
Why this work
India has been caught by surprise before. Not because the warning didn't exist, but because it existed somewhere in the system and nobody put it together in time. People died because of that gap. We're closing it.
We're building the intelligence system that helps India see the next attack coming before it happens, and helps it win the next war before the first shot is fired. Not with better sensors; India already collects enough. With the ability to actually use what it collects, at the speed the threat moves.
This doesn't get built by a foreign company, and it doesn't get built for a demo. It gets built by people who decided this mattered enough to build it here, for real, before it's needed. If we do this right, the payoff is a warning that gets acted on in time, and a war that's already won in preparation before it's fought at all.
What you'd work on
- Identity across scripts: A name in Devanagari, Roman and a third script has no canonical form. Edit distance assumes a shared alphabet. Phonetic methods were built around English. Multilingual embeddings will cheerfully merge two distinct people who share a name component, the exact failure that ends you. What replaces the standard toolkit is open.
- Decisions under asymmetric cost: A wrong merge is far more expensive than a missed one, and most of the published literature optimises a metric that assumes the two costs are equal. You'd set thresholds against that asymmetry with very few labels.
- Extraction from data that resists it: Mixed-language reports,
OCR'd handwriting, no standard format. Turning that into entities you'd stake an assessment on, with hallucination detection you actually trust, is its own problem.
- A graph that corrects itself: It grows with every document and will make mistakes. A merge made today gets revisited tomorrow, and everything built on a bad one has to be findable when it's undone. Full provenance, no quiet corruption.
- Finding what nobody asked for: Structural patterns across unrelated entities, anomalies against learned baselines, signals surfaced before anyone poses the question. Not RAG, not search. Very little exists to copy.
Why it's hard At this scale all-pairs comparison is out, and standard candidate generation assumes comparable surface forms, the exact assumption that breaks first here. Rare-event base rates are punishing, which makes overall accuracy meaningless and puts all the weight on the top of a ranked list.
Who this is for
Given an ambiguous problem, you decompose it, name the tradeoffs, and design, rather than looking for a tutorial. You understand embeddings and retrieval mechanically. You've worked on real, messy data, and you know a system that's right 95% of the time can be useless if the other 5% is catastrophic.
Direct experience with entity resolution, knowledge graphs, or multilingual NLP is a solid signal, not a prerequisite. Linguistics, IR, database theory and formal methods backgrounds all interest us. How you think matters more than what you've built.
Not this role: RAG over a document store, or maintaining an ontology someone else designed.
To apply
Send a resume, plus up to two pages:
You need to decide whether two records refer to the same real-world entity. You have no labelled pairs, and a false match is far more costly than a missed one. Describe how you'd build it, how you'd set the threshold with no labels, and what evidence would convince a sceptic that it works.
We don't have a model answer. We're reading how you reason with no paper to cite. Email:
[email protected]
📌 AI Research Engineer: Entity Resolution & Knowledge Graphs (Bengaluru)
🏢 Auric AI Labs
📍 Bengaluru