Evals for frontier research agents
We build custom evals for product teams, as a service. To build these eval sets, we source domain experts and have them turn their work into original, real world evals.
For example, our research engineering experts have create tasks involving optimizing an algorithm, deploying a model, or running experiments to solve a novel problem. A grader scores the agent’s performance, and these scores serve as signals during evaluations.