Everyone else here builds the system. This role exists to prove it's wrong.
You'd be measured on how much damage you do to our own work.
Why this work
India has been caught by surprise before. Not because the warning didn't exist, but because it existed somewhere in the system and nobody put it together in time. People died because of that gap. We're closing it.
We're building the intelligence system that helps India see the next attack coming before it happens, and helps it win the next war before the first shot is fired. Not with better sensors; India already collects enough. With the ability to actually use what it collects, at the speed the threat moves.
This doesn't get built by a foreign company, and it doesn't get built for a demo. It gets built by people who decided this mattered enough to build it here, for real, before it's needed. If we do this right, the payoff is a warning that gets acted on in time, and a war that's already won in preparation before it's fought at all.
What that means for this role The characteristic failure of an agentic reasoning system is output that's fluent, well-cited, internally coherent, and wrong. It's hardest to catch exactly where it matters most, and it gets more persuasive as the models improve. In consumer software that's a bad quarter. Here it lands on people who had no say in it.
The institutions we work with are deeply sceptical of automated judgment, and they're right to be. The only real answer is a dedicated adversarial function, independent of the teams it evaluates. A critic sharing context with its target inherits the target's blind spots. It's the highest-value thing we'll build and the easiest to quietly cut. We're putting it in the job post so you can hold us to it.
If you think AI in defence should be built carefully or not at all,
this is the role that makes carefully possible, and keeps it honest.
A note on the obvious pitch The tempting framing is self-play: agents run ten thousand scenarios overnight and find strategies no human would. We've retired it. AlphaGo had a perfect simulator, a free reward signal, stationary rules, symmetric self-play. This domain has none of the four: it's partially observable, the adversary adapts to you specifically, and outcomes are sometimes unobservable for years.
So the real question is what replaces the reward when the reward is unobservable, and what a self-play equilibrium means when the simulator is itself a hypothesis. If your instinct reading the AlphaGo line was that it breaks, that's the instinct this role runs on.
What you'd work on
- Adversaries worth the name: Not a scripted opponent, not a critique template. Something that finds the alternative explanation, isolates the assumption doing the most work and tests it, and notices the evidence that should exist if a conclusion were true and doesn't.
- Attacking our own system: Poisoning, injection embedded in material it will read, saturation, probing. An adversary who learns what your system reacts to has learned more than the data itself would give up. Then building the defences and proving they hold.
- Knowing whether any of it works: When nobody reviews every output individually, correctness stops being answerable by inspection. Held-out ground truth, scoring by question class,
variance across runs, drift detection. We think this is a bigger engineering artifact than the agents themselves. Most teams build it last.
- Whether we're making humans worse: People shown a fluent machine assessment anchor on it, and their own reasoning measurably degrades. Raising apparent throughput while quietly degrading judgment makes the institution worse while every dashboard improves. Designing the trials that catch this is part of the job.
The deepest problem here: generating objections is trivial, models do it endlessly. Telling a critique that finds a real flaw from one that's merely well-formed is not, and as far as we can tell it's unsolved. Who this is for An unusual role for an unusual person. Constitutionally sceptical, rigorous rather than reflexive about it. More excited by an experiment that disproves something than a demo that impresses someone. Able to argue for a flaw against people who don't want to hear it, including us.
Practical backgrounds: adversarial ML, security research, multi-agent systems, RL, game theory, forecasting and calibration, causal inference, experimental design. Rarer and valuable: real grounding in statistics or philosophy of science alongside the engineering. No defence background needed. No seniority needed.
To apply
Send a resume, plus the strongest argument you can make that we're wrong.
Everything you need is above: an agentic reasoning system over messy multi-source data, no fine-tuning, self-hosted models, humans holding judgment at the end of every chain. Attack it. Find the assumption we're leaning on hardest and break it.
Two pages, maximum. The best submission gets a conversation regardless of anything else in the file.
Email:
[email protected]
📌 AI Research Engineer: Adversarial Systems (Bengaluru)
🏢 Auric AI Labs
📍 Bengaluru