Location: Delhi NCR
Experiwnce - 4+ years
Core Mandate: SFT, preference optimization, reasoning RL, reward modelling, verifier environments.
Responsibilities:
Drive model alignment and capability enhancement through Supervised Fine-Tuning (SFT).
Implement and scale preference optimization techniques (e.g., RLHF, DPO) and advanced reasoning RL.
Train robust reward models and engineer complex verifier settings to support iterative capability scaling.
Requirements:
4+ years of applied experience in deep learning, specifically focusing on Reinforcement Learning, alignment methodologies, or post-training foundation models.
📌 Senior Research Engineer – Post Training & Reinforcement Learning Delhi
🏢 Eternity Quests
📍 Delhi
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.