12 Aug
|
Dusker AI
|
Jaipur
About Dusker AI
Dusker AI specializes in benchmarking and evaluating AI agents to help organizations understand real-world performance. Using expert-driven frameworks, we assess AI systems across reasoning, reliability, adaptability, and safety to ensure they are truly production-ready. From conversational AI to autonomous agentic systems, we build cutting-edge evaluation frameworks that enable organizations to develop trustworthy, high-performing AI solutions.
Role Overview
We are seeking an experienced RLHF / Post-Training Engineer to build the preference data pipelines, reward models, and alignment training loops that turn raw model capability into dependable behaviour. You will own supervised fine-tuning and preference optimization workflows end to end, from response sampling and reward modeling through training runs and post-training evaluation. You will work alongside evaluation scientists, domain experts, and infrastructure engineers to close the loop between measured model weaknesses and the training interventions that actually fix them.
This role is ideal for someone passionate about preference learning, reward modeling, alignment methodology, and rigorous measurement of post-training gains.
Key Responsibilities
- Design and run supervised fine-tuning and preference optimization pipelines using methods such as PPO, DPO, and GRPO.
- Build reward models and rubric-based graders that convert expert human judgment into reliable training signal.
- Architect preference data collection workflows covering response sampling, pairwise comparison design, and adjudication of disagreement.
- Develop reinforcement learning loops with verifiable rewards for reasoning, tool calling, and instruction following.
- Conduct systematic ablations to isolate which data mixtures, objectives, and hyperparameters genuinely drive gains.
- Collaborate with evaluation scientists to measure post-training effects on reasoning, safety,
refusal behaviour, and regression risk.
- Optimize distributed training throughput and cost using LoRA, QLoRA, mixed precision, and efficient checkpointing.
- Guide annotation leads on rater calibration, quality control, and early detection of reward hacking in collected data.
- Document training recipes, data lineage, and experiment results so that every run is independently reproducible.
Required Qualifications
- Bachelor's or Master's degree in Computer Science, Machine Learning, Statistics, or a related quantitative field.
- 4 or more years of machine learning engineering experience, including at least 2 years working directly with large language models.
- Robust Python skills and deep hands-on use of PyTorch alongside Hugging Face Transformers, TRL, PEFT, and DeepSpeed or Accelerate.
- Practical experience with post-training methods such as supervised fine-tuning, RLHF, DPO, or comparable preference optimization techniques.
- Demonstrated experience training models across multiple GPUs, including distributed strategies and memory optimization.
- Working knowledge of reward modeling and its failure modes, including reward hacking, over-optimization, and annotator bias.
- Familiarity with evaluation methodology, including held-out benchmarks, LLM-as-judge harnesses, and significance testing.
- Excellent written communication and the discipline to report results honestly, including the negative ones.
Preferred Qualifications
- Experience with reinforcement learning from verifiable rewards or reasoning-focused post-training.
- Background in MLOps tooling such as Docker, Kubernetes, CI/CD, and experiment tracking with Weights and Biases or MLflow.
- Exposure to large-scale training on AWS, Azure, or GCP, including spot capacity and fault-tolerant checkpointing.
- Open-source contributions to training, alignment, or evaluation libraries.
- Publications in alignment, preference learning, or model evaluation.
📌 Training Engineer (Jaipur)
🏢 Dusker AI
📍 Jaipur