13 Sep
|
jurisphere.ai
|
Bengaluru
13 Sep
jurisphere.ai
Bengaluru
Research Engineer, Post-Training
Location:
Bengaluru, India
Type:
Full Time
Team:
Engineering / AI Research
Why Jurisphere
Jurisphere is building AI systems that can do reliable legal work.
We recently raised a $2.2M seed round led by Info Edge Ventures
, with participation from Flourish Ventures, Antler, and 8i Ventures. Today, Jurisphere is used by 500+ organisations and 20,000+ legal professionals
, including ICICI Bank, Unilever, Philips, Tata Capital, and CMS IndusLaw.
We are a small team working at the intersection of frontier AI, legal reasoning, and high-stakes professional workflows
.
The Role
We’re hiring a Research Engineer, Post-Training to help make our models meaningfully better at legal work.
You will work across model training, datasets, evaluations, agents, and expert feedback — taking problems from an observed model failure through to an experiment, training signal, and measurable improvement.
This is a hands-on research engineering role for someone comfortable working with open-weight models and ambiguous research problems.
What You’ll Do
- Run post-training experiments across SFT, preference optimization, RLHF/RLAIF, reward modelling, distillation, and related techniques .
- Build training datasets from expert feedback, model traces, legal workflows, and synthetic data .
- Design evaluations and reward systems for legal reasoning, research, drafting, document analysis, and citation accuracy.
- Study model and agent behaviour to identify systematic failure modes.
- Build agent environments involving retrieval, tools, subagents, validation loops, and long-horizon workflows .
- Turn research findings into better training data, model recipes, evaluations, and production systems .
- Work closely with engineers and legal experts to translate expert judgment into scalable systems.
What You Have
- Hands-on experience training or post-training LLMs , particularly open-weight models.
- Strong Python and ML engineering skills.
- Experience with some combination of SFT, preference learning, RL, reward modelling, distillation, agents, or evaluations .
- Strong experimental judgment: you can formulate hypotheses, run controlled experiments, analyse model behaviour, and interpret ambiguous results.
- Ability to independently own technically difficult, loosely defined problems.
Nice to have: distributed training or GPU experience, agent systems, evaluation infrastructure, research publications, open-source contributions, or experience in high-stakes domains such as law, finance, or healthcare.
Why This Role
Post-training at Jurisphere is the feedback loop between models, real legal work, expert judgment, and engineering
.
You’ll work across the full loop: understand why a model failed, build an evaluation that captures it, create a training signal to improve it, and measure whether the next model is actually better.
The goal: build models that are meaningfully better at doing legal work.
📌 Research Engineer, Post-Training (Bengaluru)
🏢 jurisphere.ai
📍 Bengaluru