nWe're training the next generation of Vedika models.
nThis role sits directly inside RL and post-training research: designing how the model learns after pretraining, how it improves from interaction, how it handles long-horizon tasks, and how we push capability beyond standard instruction tuning.
nYou'll work on:
n
n
- RL training for next-generation Vedika models
n
- GRPO, PPO, DPO and newer post-training methods
n
- Reward models, process rewards and verifiers
n
- Long-horizon reasoning and agent trajectories
n
- Tool-use and computer-use reinforcement
n
- Self-improvement and synthetic training loops
n
- Multi-turn behaviour and memory training
n
- Failure mining from model trajectories
n
- Evaluation systems for reasoning, autonomy and reliability
n
- Research experiments that can become part of the next model generation
n
nCompensation:
n₹2.6 LPA fixed
n₹3.6 LPA CTC
nWork mode:
Fully remote
nYou'll get:
n
n
- Mac for development
n
- Claude
n
- Codex
n
- Reliable compute and research infrastructure
n
- ₹10L–₹50L+ yearly AI/token spend available across the team and experiments
n
nThis is not a role for someone whose idea of model work ends at prompting or basic fine-tuning.
nWe want someone who can understand a training run, break it, diagnose it, redesign it and make the next model measurably better.
nStrong PyTorch, RL fundamentals, post-training, distributed training and hands-on experimentation matter far more than credentials.
nRole: Benjamin RL
nVedika — Next Generation Models
nSend your work, experiments, papers, repos or anything you trained that genuinely got better.
📌 Benjamin RLHF (Pune)
🏢 Vedika API
📍 Pune
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.