Job Description
Weekly Hours: 5 – 25+ hours per week (Flexible & Asynchronous)
what suits you.
• Start: Immediately. The client has sample work that needs turning around this week, with a
substantially larger programme expected to follow.
About the Role
We are seeking exceptional Software Engineers to join our team as AI Benchmark Task Authors. In this role, you will not simply write standard feature code; you will design, engineer, and author complex, production-grade software evaluation environments used to benchmark frontier Artificial Intelligence models (such as GPT-5, Claude, and Gemini).
You will build realistic multi-module code repositories, author ambiguous real-world issue specifications, and write hidden automated test suites designed to stress-test model reasoning, edge-case handling, and architectural correctness.
Key Responsibilities
- Repository Fixture Engineering: Design and build self-contained, dependency-light, multi-module software codebases (1,000–2,000 lines of code) in languages like Python, TypeScript, C++, Java, or Go.
- Specification Authoring: Write explicit issue descriptions framing problems or feature requirements from an end-user perspective—without revealing the exact code paths or files that need modification.
- Hidden Test Suite Authoring: Construct comprehensive hidden unit and integration test suites (pytest, Jest, etc.) that evaluate exact functional correctness, boundary conditions, edge cases, and refusal honesty.
- Gold & Partial Solution Authoring: Implement 100% correct Gold reference solutions alongside plausible, imperfect partial credit implementations to ensure test suites measure a gradient of model understanding.
- Quality & Contamination Assurance: Ensure all authored fixtures are clean, leak-free, and deterministic with zero network or wall-clock dependencies.
Key Requirements & Qualifications
- Core Backend Proficiency: Strong hands-on development experience in at least one primary backend language (Python,
📌 AI Trainer (India)
🏢 Devfixr
📍 India