14 Sep
|
Vedika API
|
India
Hiring: Benjamin RL We’re training the next generation of Vedika models. This role sits directly inside RL and post-training research: designing how the model learns after pretraining, how it improves from interaction, how it handles long-horizon tasks, and how we push capability beyond standard instruction tuning. You’ll work on:
- RL training for next-generation Vedika models - GRPO, PPO, DPO and newer post-training methods - Reward models, process rewards and verifiers - Long-horizon reasoning and agent trajectories - Tool-use and computer-use reinforcement - Self-improvement and synthetic training loops - Multi-turn behaviour and memory training - Failure mining from model trajectories - Evaluation systems for reasoning, autonomy and reliability - Research experiments that can become part of the next model generation Compensation: ₹2.6 LPA fixed ₹3.6 LPA CTC Work mode:
Fully remote You’ll get:
- Mac for development - Claude - Codex - Responsible compute and research infrastructure - ₹10L–₹50L+ yearly AI/token spend available across the team and experiments This is not a role for someone whose idea of model work ends at prompting or basic fine-tuning. We want someone who can understand a training run, break it, diagnose it, redesign it and make the next model measurably better.
Strong
PyTorch, RL fundamentals, post-training, distributed training and hands-on experimentation matter far more than credentials.
Role: Benjamin RL Vedika — Next Generation Models Send your work, experiments, papers, repos or anything you trained that genuinely got better.
📌 Benjamin Rlhf (India)
🏢 Vedika API
📍 India