Methodology

How a pull request becomes a task

Pipeline6 stepsmine → validate → specify → seal → tier → verify
Verification3×every task graded three times in its shipped image
Protocol100 stepsmedium effort · no network · graded as-is at the limit
Corrected in verification60 taskspatch or test list fixed before release · 3 excluded
The pipelinesix steps from a merged pull request to a shipped task
  1. 1 · mine

    Find the bug

    Merged pull requests that fix concurrency and async-lifecycle bugs in permissively licensed Python projects.

  2. 2 · validate

    Prove it reproduces

    A 2×2 matrix run three times: the new tests fail without the fix, pass with it, nothing else breaks. Flaky tasks are dropped.

  3. 3 · specify

    Write the bug report

    The failure and the interface the tests rely on, with no hint of the fix. An auditor checks it covers every hidden test.

  4. 4 · seal

    Probe for leaks

    Each sandbox is checked for a single commit, no remotes, blocked network and no trace of hidden test names.

  5. 5 · tier

    Measure difficulty

    Calibration runs with low-cost models sort tasks into clean, split and hard. Composites combine several fixes.

  6. 6 · verify

    Check what ships

    Each image is built and graded three times with the reference fix (must score 1) and an empty submission (must score 0).

What a task looks likea race condition, and what the agent is given
time → Task A read 4 await write 5 count 4 5 (expected 6) Task B read 4 write 5 Both tasks read 4 before either writes. One increment is lost.
The shape of bug these tasks are made of: the code is correct in any single ordering and wrong in one interleaving.
seesA bug report and the repository at the commit before the fix.
gitOne commit. Later commits, branches, tags, remotes and the reflog are gone.
networkNone. The sandbox reaches only the model endpoint.
testsHidden until grading, then run in a separate, clean verifier container.
rewardBinary: all fail-to-pass tests pass and no pass-to-pass test regresses.
composites74 tasks combine several upstream fixes to one project, so the complete solution never existed as one commit.
Release checkgraded the way it will be run
  • Every task rebuilt from its published files and graded end to end, in the exact container that ships.
  • Test ids matched against what pytest actually collects.
  • A fail-to-pass test must fail without the fix and pass with it; a pass-to-pass test must pass both ways; in every one of three runs.
  • 60 tasks ship with a patch or test list corrected by that check.
  • Each image is scanned for any copy of the reference fix outside the repository.

Excluded after verification:

  • ACCORD-NWP/tactus #182: race reproduces intermittently: the empty submission scored 1 in release_check job 2168 and 0,0,0 in job 2194, so a no-op agent would pass about 1 run in 4
  • 2 private task(s), for reasons of the same kinds (a race that does not always reproduce, a flaky regression test).
How the protocol comparesDeepSWE, SWE-bench, Terminal-Bench
BenchmarkStep limitCost capTimeoutRuns per taskAt the limit
SWE-Race100 stepsnone600 s per command2–6, mean pass@1graded as-is
DeepSWE100 stepsnonenot stated16, mean pass@1graded as-is
SWE-bench Verified (bash only)250 steps$360 s per command1empty patch
SWE-bench Pro250 turnsnone60 s read1partial diff graded
Terminal-Bench 2.0nonenone10–200 min per task≥5not graded, 0

From each benchmark's published documentation, October 2026. SWE-Race follows DeepSWE's protocol and adds repeated attempts, confidence intervals and cost.

Contaminationolder tasks against newer tasks, by size of the fix

Tasks come from merged pull requests, so each reference fix is public. Resolve rate by size of the reference fix, tasks merged before 2026 against tasks merged in 2026:

≤10 lines
99.2% · n=13
90.0% · n=19
11–50
87.1% · n=27
84.9% · n=44
51–200
82.9% · n=21
68.2% · n=44
>200
60.0% · n=7
55.9% · n=13
before 20262026

Older tasks score higher in most size bands, which is what memorisation would look like, but the gap is not statistically significant at this sample size. Each task records its merge date so any result can be split by era.