What if your years of research expertise could directly shape the trajectory of AI development? We're looking for elite machine learning specialists to design the kinds of problems that push state-of-the-art AI systems to their absolute limits — and help the field understand precisely where they break down.
This is a rare, high-impact prospect to put your hard-earned domain knowledge to work at the frontier of AI evaluation and safety. Your contributions won't collect dust in a journal — they'll actively influence how the next generation of AI models is built, tested, and improved.
Design complex, original machine learning problems rooted in your specialized domain of expertise
Craft evaluation tasks that demand advanced knowledge far beyond standard ML pipelines
Draw from your own research to create challenges that would stump even the most capable LLMs
Define precise problem statements,
gold-standard solutions, and robust evaluation criteria
Assess AI-generated ML solutions for correctness, creativity, and methodological rigor
Document expected failure modes, difficulty levels, and the domain knowledge required to solve each problem
Collaborate asynchronously with a global network of top researchers and engineers
Who You Are
You hold a graduate degree (MS or PhD) in a scientific or technical field that intersects with machine learning
You have strong working knowledge of ML fundamentals — model selection, feature engineering, evaluation metrics, and pipeline design
You're deeply embedded in active research problems in your field and know where the hard edges are
You can identify precisely where general ML knowledge falls short and specialized domain expertise becomes essential
You have experience publishing or conducting original research (highly valued)
You communicate complex idea