14 Sep
|
Vedika API
|
Maharashtra
14 Sep
Vedika API
Maharashtra
Hiring: Benjamin RLWe’re training the next generation of Vedika models.This role sits directly inside RL and post-training research: designing how the model learns after pretraining, how it improves from interaction, how it handles long-horizon tasks, and how we push capability beyond standard instruction tuning.You’ll work on:
RL training for next-generation Vedika models
GRPO, PPO, DPO and newer post-training methods
Reward models, process rewards and verifiers
Long-horizon reasoning and agent trajectories
Tool-use and computer-use reinforcement
Self-improvement and synthetic training loops
Multi-turn behaviour and memory training
Failure mining from model trajectories
Evaluation systems for reasoning, autonomy and reliability
Research experiments that can become part of the next model generationCompensation:₹2.6 LPA fixed₹3.6 LPA CTCWork mode:
Fully remoteYou’ll get:
Mac for development
Claude
Codex
Responsible compute and research infrastructure
₹10L–₹50L+ yearly AI/token spend available across the team and experimentsThis is not a role for someone whose idea of model work ends at prompting or basic fine-tuning.We want someone who can understand a training run, break it, diagnose it, redesign it and make the next model measurably better.Robust PyTorch, RL fundamentals, post-training, distributed training and hands-on experimentation matter far more than credentials.Role: Benjamin RLVedika — Next Generation ModelsSend your work, experiments, papers, repos or anything you trained that genuinely got better.
📌 Benjamin Rlhf Maharashtra
🏢 Vedika API
📍 Maharashtra