We build the evaluation layer that validates AI agents before they reach customers — an automated system that scores across large volumes of agent traces.
You'll own core parts of that platform: the pipeline that runs traces through model-based judges at scale, and the scoring logic that turns raw output into results teams can act on.
What you'll do
- Design evaluation methodologies and benchmarks for agent reasoning, planning, tool use, reliability, and safety — across LLM-as-a-Judge, trajectory-based, and human evaluation
- Take problems from research question to prototype to shipped feature, owning them end to end
- Build and harden the pipelines and scoring logic behind customer-facing evaluation
- Curate synthetic and real-world datasets; measure the evaluator itself for consistency and agreement with human labels