Measuring AI systems on the work they are used for
Dedicated benchmarks and evaluations for models, model routers and agents: built from real work, graded by execution, reported with uncertainty.
Benchmarks
Public task families with a private held-out split, like SWE-Race.
Evaluations
Models, routers, harnesses and prompts, on the same tasks, so the gain from each is measured, not assumed.
How we work
Real tasks, executable grading, sealed sandboxes, verified tests, published failures. Open to research cooperation: get in touch →
What we evaluatemodels, routers, agents, prompts
Models
Capability on a defined task family: resolve rate with confidence intervals, pass@k across repeated attempts, cost per solved task.
Routers
Systems that choose a model per request. We run the router and every model it can route to on the same tasks and compare quality, cost and latency.
Agents and harnesses
The same model behaves differently under different scaffolds, tools and step limits. We hold the tasks and the container fixed and vary the harness.
Prompts and workflows
Regression suites for production prompts and multi-step workflows, so a change can be judged before it ships. This is the evaluation layer behind Evaligo.
How we build a benchmarksix steps, the same for every task family
- 1 · source
Start from real work
Merged fixes in active open-source projects, production workflows, real user requests. Not synthetic puzzles.
- 2 · execute
Make grading executable
A reproducible environment and a hidden check that runs, such as the project's own tests.
- 3 · specify
Describe the problem, not the answer
Each prompt states the failure and the interface the checks rely on, audited against the hidden tests.
- 4 · seal
Close the shortcuts
No network, no future git history, no hidden test names. Each sandbox is probed before use.
- 5 · verify
Test the test
The reference solution must pass and an empty submission must fail, repeatedly, in the container that ships.
- 6 · hold out
Keep a private split
Part of every benchmark is never published, so a public/private gap exposes overfitting.
Principlesfour rules we hold ourselves to
- Report uncertainty. Every score has its interval and the attempts behind it; a difference inside the noise is a tie.
- Measure contamination. When tasks come from public history, we test whether older tasks are easier and publish it.
- Same conditions for everyone. One harness, one step limit, one environment per task; we run every model ourselves.
- Show the failures. Excluded tasks, corrected tests and known limits are published with the methodology.
Need an evaluation built for your use case?
Private benchmarks for teams choosing models, tuning routers or shipping agents.