14 Sep
|
Vedika API
|
Mumbai
Job DescriptionHiring: Benjamin RL NnWe’re training the next generation of Vedika models. NnThis role sits directly inside RL and post-training research: designing how the model learns after pretraining, how it improves from interaction, how it handles long-horizon tasks, and how we push capability beyond standard instruction tuning. NnYou’ll workon: N
- RL training for next-generation Vedika models N
- GRPO, PPO, DPO and newer post-training methods N
- Reward models, process rewards and verifiers N
- Long-horizon reasoning and agent trajectories N
- Tool-use and computer-use reinforcement N
- Self-improvement and synthetic training loops N
- Multi-turn behaviour and memory training N
- Failure mining from model trajectories N
- Evaluation systems for reasoning, autonomy and reliability N
- Research experiments that can become part of the next model generation NnCompensation: N₹2.6 LPA fixedn₹3.6 LPA CTC NnWork mode:
Fully remote NnYou’ll get: N
- Mac for developmentn
- Claude N
- Codexn
- Seriouscompute and research infrastructure N
- ₹10L–₹50L+ yearly AI/token spend available across the team and experiments NnThis is nota role for someone whose idea of model work ends at prompting or basic fine-tuning. NnWe want someone who can understand a training run, break it, diagnose it, redesign it and make the next model measurably better. NnStrong PyTorch, RL fundamentals, post-training, distributed training and hands-on experimentation matter far more than credentials. NnRole: Benjamin RL NVedika — Next Generation Models NnSend your work, experiments, papers, repos or anything you trained that genuinely got better.
📌 Benjamin Rlhf (Mumbai)
🏢 Vedika API
📍 Mumbai