We are hiring Reward Validation Specialists to contribute to advanced AI training and evaluation projects. This role is ideal for PhD-qualified researchers with expertise in reinforcement learning, optimization, and Python programming who are passionate about improving the reliability and accuracy of AI evaluation systems.
Role Overview In this role, you will help build and validate automated grading systems for AI model evaluation pipelines. Rather than assessing only the final output, you will evaluate and verify step-by-step reward logic to ensure each grading component accurately measures the intended behavior. You will also develop and implement automated graders in Python that can scale across large evaluation workflows while maintaining consistency, accuracy, and alignment with task requirements. Key Responsibilities
Validate step-level reward logic to ensure it accurately measures the intended model behavior.
Verify that grading criteria align with task instructions and evaluation objectives.
Ensure models have the required context, files, and information needed to complete each evaluated step.
Develop, implement, and test automated graders using Python.
Analyze and improve reward functions to support scalable and reliable AI evaluation pipelines.
Collaborate with internal teams to refine evaluation methodologies and maintain grading consistency Mandatory Requirements
PhD (or PhD candidate nearing completion) in Computer Science, Machine Learning, Artificial Intelligence,
Electrical Engineering, Applied Mathematics, Robotics, or another quantitative discipline.
Strong understanding of reinforcement learning concepts, including Markov Decision Processes (MDPs), value functions, credit assignment, and reward shaping.
Solid foundation in optimization techniques, including gradient-based methods and convex/non-convex optimization.
Strong Python programming skills with experience writing, testing, and debugging code.
Ability to quickly learn and independently apply new evaluation methodologies.
Preferred
Qualifications
Experience with control theory, energetic programming, optimal control, or stability analysis.
Familiarity with RLHF (Reinforcement Learning from Human Feedback), process reward models, or AI evaluation frameworks.
Hands-on experience with NumPy, PyTorch, or other machine learning libraries.
Background in machine learning, reinforcement learning, robotics, control systems, or applied mathematics.
Experience developing evaluation pipelines, automated testing frameworks, or model validation tools. ~Please be aware of recruitment scams involving individuals or organizations falsely claiming to represent employers. Innodata will never ask for any form of payment during any stage of the recruitment, onboarding or employment process. If you believe you’ve been targeted by a recruitment scam, please immediately report it to
[email protected]
📌 Domain Expert - Mathematics (India)
🏢 Innodata
📍 India