29 Aug
|
Innodata India
|
Noida
29 Aug
Innodata India
Noida
ROLE OVERVIEW & OBJECTIVE:
We are expanding our core evaluation and are looking to hire experienced AI Evaluation and Reinforcement Learning from Human
Feedback (RLHF) Specialists into our high-complexity alignment pipeline. This cohort directly powers reward model training, preference modeling, complex multi-turn reasoning, and critique-and-rework loops for top-tier frontier models. We are sourcing experienced practitioners from leading AI data platforms who require zero ramp-up time, demonstrate rigorous adherence to energetic rubrics, and deliver ironclad, defensible evaluation rationales.
KEY RESPONSIBILITIES & CORE WORKFLOWS:
- High-Precision Preference Data Generation: Author high-fidelity preference datasets and rank multi-turn model completions to
optimize reward model loss functions.
- Side-by-Side Evaluations: Execute deep SxS comparative evaluations across factuality, helpfulness, harmlessness, logical
coherence, and instruction-following.
- Improve Model output: Rewrite sub-optimal model outputs and break down complex chain-of-thought failures into structured,
corrective learning signals.
- Multi-Step Reasoning & Agentic Auditing: Evaluate step-by-step mathematical, logical, and tool-use trajectories generated by
autonomous AI agents.
- Golden Sets & Worked Examples:
Construct reference 'golden responses' and edge-case worked examples to train and calibrate
downstream evaluation cohorts.
- Inter-Rater Reliability (IRR) Leadership: Participate in daily calibration huddles, driving consensus on ambiguous prompts and rapid
guideline iterations.
Mandatory Requirements:
- Direct AI Platform Experience: 2+ years of hands-on data
annotation, model evaluation, or RLHF preference scoring on recognized platforms
- Complex Task Track Record: Demonstrable mastery in multidimensional rubric scoring, and chain-of-thought verification.
- Written Communication: English proficiency with proven
discipline in articulating clear, fact-backed, objective rationales under tight SLAs.
- Operational Agility: High tolerance for rapidly evolving project
guidelines, shifting evaluation criteria, and fast-paced delivery sprints.
Preferred Qualifications:
- Domain Specialization: Advanced educational background in
STEM, Law, Finance, Medicine, or Humanities.
- Technical Literacy: Basic Python scripting, prompt engineering
knowledge, or familiarity with loss functions and reward modeling concepts.
- Leadership Experience: Prior experience mentoring junior raters,
conducting secondary QA audits, or serving as a task calibration lead.
📌 Senior AI Evaluation & RLHF Specialist (Noida)
🏢 Innodata India
📍 Noida