About Evaligo Labs

Measuring AI systems on the work they are used for

Dedicated benchmarks and evaluations for models, model routers and agents: built from real work, graded by execution, reported with uncertainty.

Benchmarks

Public task families with a private held-out split, like SWE-Race.

Evaluations

Models, routers, harnesses and prompts, on the same tasks, so the gain from each is measured, not assumed.

How we work

Real tasks, executable grading, sealed sandboxes, verified tests, published failures. Open to research cooperation: get in touch →

What we evaluatemodels, routers, agents, prompts

Models

Capability on a defined task family: resolve rate with confidence intervals, pass@k across repeated attempts, cost per solved task.

Routers

Systems that choose a model per request. We run the router and every model it can route to on the same tasks and compare quality, cost and latency.

Agents and harnesses

The same model behaves differently under different scaffolds, tools and step limits. We hold the tasks and the container fixed and vary the harness.

Prompts and workflows

Regression suites for production prompts and multi-step workflows, so a change can be judged before it ships. This is the evaluation layer behind Evaligo.

How we build a benchmarksix steps, the same for every task family
  1. 1 · source

    Start from real work

    Merged fixes in active open-source projects, production workflows, real user requests. Not synthetic puzzles.

  2. 2 · execute

    Make grading executable

    A reproducible environment and a hidden check that runs, such as the project's own tests.

  3. 3 · specify

    Describe the problem, not the answer

    Each prompt states the failure and the interface the checks rely on, audited against the hidden tests.

  4. 4 · seal

    Close the shortcuts

    No network, no future git history, no hidden test names. Each sandbox is probed before use.

  5. 5 · verify

    Test the test

    The reference solution must pass and an empty submission must fail, repeatedly, in the container that ships.

  6. 6 · hold out

    Keep a private split

    Part of every benchmark is never published, so a public/private gap exposes overfitting.

Principlesfour rules we hold ourselves to
  • Report uncertainty. Every score has its interval and the attempts behind it; a difference inside the noise is a tie.
  • Measure contamination. When tasks come from public history, we test whether older tasks are easier and publish it.
  • Same conditions for everyone. One harness, one step limit, one environment per task; we run every model ourselves.
  • Show the failures. Excluded tasks, corrected tests and known limits are published with the methodology.

Need an evaluation built for your use case?

Private benchmarks for teams choosing models, tuning routers or shipping agents.

Tell us what you need