Benjamin Rlhf Pune

Benjamin Rlhf Pune

13 Sep
|
Vedika API
|
Pune

13 Sep

Vedika API

Pune

Hiring: Benjamin RL

We’re training the next generation of Vedika models.

This role sits directly inside RL and post-training research: designing how the model learns after pretraining, how it improves from interaction, how it handles long-horizon tasks, and how we push capability beyond standard instruction tuning.

You’ll work on
RL training for next-generation Vedika models
GRPO, PPO, DPO and newer post-training methods
Reward models, process rewards and verifiers
Long-horizon reasoning and agent trajectories
Tool-use and computer-use reinforcement
Self-improvement and synthetic training loops
Multi-turn behaviour and memory training
Failure mining from model trajectories
Evaluation systems for reasoning, autonomy and reliability
Research experiments that can become part of the next model generation

Compensation

₹2.6 LPA fixed

₹3.6 LPA CTC

Work mode: Fully remote





You’ll get
Mac for development
Claude
Codex
Reliable compute and research infrastructure
₹10L–₹50L+ yearly AI/token spend available across the team and experiments

This is not a role for someone whose idea of model work ends at prompting or basic fine-tuning.

We want someone who can understand a training run, break it, diagnose it, redesign it and make the next model measurably better.

Solid PyTorch, RL fundamentals, post-training, distributed training and hands-on experimentation matter far more than credentials.

Role: Benjamin RL

Vedika — Next Generation Models

Send your work, experiments, papers, repos or anything you trained that genuinely got better.

📌 Benjamin Rlhf Pune
🏢 Vedika API
📍 Pune

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: benjamin rlhf pune / pune