About Auric AI
We're building a reasoning system over millions of messy, multilingual intelligence documents, and it has to run entirely inside India, on infrastructure we control.
No fine-tuning. No external APIs. Self-hosted open-weight models, air-gapped. Every point of performance comes out of architecture. We're hiring one research engineer to own it.
Why this works
India has been caught by surprise before. Not because the warning didn't exist, but because it existed somewhere in the system and nobody put it together in time. People died because of that gap. We're closing it.
We're building the intelligence system that helps India see the next attack coming before it happens, and helps it win the next war before the first shot is fired. Not with better sensors; India already collects enough. With the ability to actually use what it collects, at the speed the threat moves.
This doesn't get built by a foreign company, and it doesn't get built for a demo. It gets built by people who decided this mattered enough to build it here, for real, before it's needed. If we do this right, the payoff is a warning that gets acted on in time, and a war that's already won in preparation before it's fought at all.
What you'd work on
- Retrieval that knows what it missed: A production RAG system returns a confident answer and has no idea what it failed to surface. Here that's the one failure we can't ship. What it would take for retrieval to bound its own recall is open.
- Reasoning across many hops and sources: Real questions don't resolve in one lookup. They need decomposition, retrieval that notices its own gaps, and synthesis across sources that disagree, with contradictions surfaced rather than averaged away.
- Uncertainty that survives the chain: Evidence varies wildly in reliability and precision,
and inference runs several steps deep. Confidence has to propagate explicitly, because a wrong answer delivered with false confidence is worse than no answer.
- Investigations longer than a context window: Work spanning days and hundreds of tool calls can't live in context. What gets externalised, compressed, and reconstructed is mostly unsettled, and everything else depends on it.
Why it's hard The models are fixed, so architecture is the only lever. The data resists every clean assumption: decades of documents, a dozen languages, no schema, the same entity written five different ways. Nothing gets to be a black box; a person accountable for a decision won't act on a system they can't interrogate, which rules out a lot of otherwise convenient architectures. No standard playbook exists. You'd derive it.
Who this is for
No experience requirement, no degree requirement. We're reading for one thing: given an open problem with no paper to follow, can you reason your way to an architecture, build it, and measure it honestly.
You should understand language models and retrieval mechanically, not as APIs. Why naive RAG fails on multi-hop temporal questions should be something you can explain without a blog post. You should have built something that survived real, messy data. A publication record is a real signal, not a requirement. So is a repository that does something nobody asked for.
Not this role: prompt templates, API integration, backend or UI, fine-tuning on labelled datasets.
To apply
Send a resume, plus up to two pages on this:
Describe the hardest retrieval or reasoning system you'd want to build, in any domain. What would you build first, what do you expect to break, and what measurement would tell you you were wrong?
Email
[email protected]. We read everything and reply either way.
📌 AI Research Engineer: Reasoning & Retrieval (Bengaluru)
🏢 Auric AI Labs
📍 Bengaluru