How a pull request becomes a task
The pipelinesix steps from a merged pull request to a shipped task
- 1 · mine
Find the bug
Merged pull requests that fix concurrency and async-lifecycle bugs in permissively licensed Python projects.
- 2 · validate
Prove it reproduces
A 2×2 matrix run three times: the new tests fail without the fix, pass with it, nothing else breaks. Flaky tasks are dropped.
- 3 · specify
Write the bug report
The failure and the interface the tests rely on, with no hint of the fix. An auditor checks it covers every hidden test.
- 4 · seal
Probe for leaks
Each sandbox is checked for a single commit, no remotes, blocked network and no trace of hidden test names.
- 5 · tier
Measure difficulty
Calibration runs with low-cost models sort tasks into clean, split and hard. Composites combine several fixes.
- 6 · verify
Check what ships
Each image is built and graded three times with the reference fix (must score 1) and an empty submission (must score 0).
What a task looks likea race condition, and what the agent is given
Release checkgraded the way it will be run
- Every task rebuilt from its published files and graded end to end, in the exact container that ships.
- Test ids matched against what pytest actually collects.
- A fail-to-pass test must fail without the fix and pass with it; a pass-to-pass test must pass both ways; in every one of three runs.
- 60 tasks ship with a patch or test list corrected by that check.
- Each image is scanned for any copy of the reference fix outside the repository.
Excluded after verification:
- ACCORD-NWP/tactus #182: race reproduces intermittently: the empty submission scored 1 in release_check job 2168 and 0,0,0 in job 2194, so a no-op agent would pass about 1 run in 4
- 2 private task(s), for reasons of the same kinds (a race that does not always reproduce, a flaky regression test).
How the protocol comparesDeepSWE, SWE-bench, Terminal-Bench
| Benchmark | Step limit | Cost cap | Timeout | Runs per task | At the limit |
|---|---|---|---|---|---|
| SWE-Race | 100 steps | none | 600 s per command | 2–6, mean pass@1 | graded as-is |
| DeepSWE | 100 steps | none | not stated | 16, mean pass@1 | graded as-is |
| SWE-bench Verified (bash only) | 250 steps | $3 | 60 s per command | 1 | empty patch |
| SWE-bench Pro | 250 turns | none | 60 s read | 1 | partial diff graded |
| Terminal-Bench 2.0 | none | none | 10–200 min per task | ≥5 | not graded, 0 |
From each benchmark's published documentation, October 2026. SWE-Race follows DeepSWE's protocol and adds repeated attempts, confidence intervals and cost.
Contaminationolder tasks against newer tasks, by size of the fix
Tasks come from merged pull requests, so each reference fix is public. Resolve rate by size of the reference fix, tasks merged before 2026 against tasks merged in 2026:
Older tasks score higher in most size bands, which is what memorisation would look like, but the gap is not statistically significant at this sample size. Each task records its merge date so any result can be split by era.