Contact
Tell us what you need
The full task set
Public and private halves, for teams that evaluate internally. We ask who you are so the private half stays private.
A model on the board
We run it under the same protocol as every other model and report the cost back to you.
Research cooperation
Shared task families, held-out evaluation of your training runs, joint write-ups.
A benchmark for you
Private benchmarks from your own codebases and workflows, for choosing models, routers or agents.