Coding-agent benchmark · Python · concurrency

SWE-Race

188 real concurrency bugs from 106 open-source projects, graded by each project's own hidden tests in a sealed, offline container. Half the tasks are held out.

Best pass@181.9%GLM-5.3 Flash · tied with GPT-5.6 Luna, DeepSeek V4 Flash
Hard subset49.9%GLM-5.3 Flash · 55 tasks
Cheapest solved task$0.018GPT-5.6 Luna · per solved task
Scale188 tasks106 projects · 2,006 attempts · 3 models

Leaderboard

RankModelpass@1Hard$ / solved taskAttemptsdetails
Tap a row for cost per solved task, attempts and more.Ranks are shared when 95% intervals overlap.All models: BenchGen agent, medium effort, 100 steps, offline.v0.1 · updated 3 Oct 2026 · changelogRequest the full task set →
How to read this table
pass@1
Share of attempts where every hidden fail-to-pass test passes and no pass-to-pass test regresses, averaged per task.
Hard
pass@1 on the 55 tasks the calibration models usually missed; the subset that separates models.
$ / solved task
Average model API cost per attempt divided by pass@1. What a success costs, not an attempt.
Interval
95% bootstrap over tasks. Shown as the thin line under each score.
Rank
Scale's rule: 1 + the number of models whose lower bound is above this model's upper bound. Overlapping intervals share a rank.
Protocol
BenchGen agent, medium reasoning effort, 100-step limit, no network. The patch present at the limit is graded as-is, as in DeepSWE. Every attempt against the shipped version of a task counts; attempts against earlier versions are discarded.
Scheduled
Claude Opus 5, GPT-5.6 Sol and other frontier models, not yet run.

Full methodology · Leak audit · results.json

What the tasks are

Real race conditions from 106 Python projects: lost updates, deadlocks, leaked cancellation. 74 combine several upstream fixes.

Browse the 95 public tasks →
How it is graded

Sealed repository at one commit, no network, hidden tests run in a clean verifier, binary reward. Every task verified three times before release.

Methodology →
Run it
git clone https://github.com/versocr/swe-race
uv tool install datacurve-pier
pier run -p swe-race/tasks --agent mini-swe-agent --model <provider/model>
Get on the board →