02 Aug
|
OpenTrain AI
|
India
02 Aug
OpenTrain AI
India
About OpenTrainOpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We help people start and grow careers teaching AI by discovering projects, building a unified profile, and applying quickly to roles that match their skills.OpenTrain connects experienced contributors and newcomers alike to meaningful AI-training work the human effort that shapes how modern AI systems behave.About AI training workAI training (data labeling, annotation, and human feedback) is the human side of building AI. Contributors annotate code, write and evaluate model outputs, and create benchmarks that teach and test models.This work is remote, versatile, often accessible without prior labeling experience, and places you on the cutting edge of how state-of-the-art AI systems are built and evaluated.The role what an AI Benchmark Engineer doesYou will design and build multi-agent benchmark tasks derived from real open-source code changes (bug fixes, migrations, refactors) and ensure those tasks run reproducibly inside containerized evaluation environments.The role centers on precise technical specifications, Python-based verification, Dockerized task execution, and decomposing complex edits across coordinated sub-agents.What youll do day-to-dayBuild multi-agent benchmark tasks grounded in real open-source code changes, including bug fixes, migrations, and refactors.Work with the Harbor evaluation framework to run and validate tasks inside Docker environments.Write clear, precise task instructions specifying file paths, function signatures, expected behavior, and constraints.Design and implement Python verification scripts to validate correctness of agent-generated code changes.Create decomposition strategies that split complex code changes across independent sub-agents.Run, debug,
and refine tasks within containers to ensure reproducibility and determinism.Evaluate task performance signals and iterate on task quality, clarity, and difficulty.RequirementsThis role requires solid practical engineering experience and familiarity with evaluation tooling and workflows.5 years of experience in Python and JavaScript development.Experience with AI coding benchmarks (for example, SWE-bench, Terminal-Bench).Strong experience reading and navigating large open-source codebases (Django, Flask, FastAPI, Node.js, or similar).Familiarity with Git workflows: pull requests, diffs, cherry-picking, and working with specific commits.Comfortable working with Docker, including writing Dockerfiles, building images, and debugging containers.Experience writing test scripts (pytest, unittest, or custom assertion-based testing).Ability to write clear, accurate, and unambiguous technical specifications.Helpful backgroundInterest in AI agentic behavior and multi-agent coordination.Familiarity with LLM evaluation and reasoning benchmarks is a plus.Location, schedule, and contract detailsThis is a remote contractor role open to contributors located in Bangladesh, Brazil, Colombia, Egypt, Ghana, India, Indonesia, Kenya, Nigeria, Turkey, or Vietnam.Time expectations: 20 hours per week, with a stated requirement of 8 hours per day and a 4-hour overlap with Pacific Standard Time (PST). Duration: 4-week contract. This contractor position does not include medical or paid leave.Who should apply and how it worksApply if you are an experienced Python/JavaScript engineer who enjoys designing reproducible evaluation tasks and working with containers and verification tooling. This role suits engineers who can read large codebases and translate changes into precise, testable tasks.To apply, create or use your OpenTrain account, complete your profile, and submit your application. OpenTrain manages hiring and contracting for this role. .
📌 AI Benchmark Engineer Software Engineering (India)
🏢 OpenTrain AI
📍 India