Coding-agent benchmark · Python · concurrency
SWE-Race
188 real concurrency bugs from 106 open-source projects, graded by each project's own hidden tests in a sealed, offline container. Half the tasks are held out.
Best pass@181.9%GLM-5.3 Flash · tied with GPT-5.6 Luna, DeepSeek V4 Flash
Hard subset49.9%GLM-5.3 Flash · 55 tasks
Cheapest solved task$0.018GPT-5.6 Luna · per solved task
Scale188 tasks106 projects · 2,006 attempts · 3 models
| Rank | Model | pass@1 | Hard | $ / solved task | Attempts | details |
|---|
Top-left is good: more tasks solved for less money. Full set.
Tap a row for cost per solved task, attempts and more.Ranks are shared when 95% intervals overlap.All models: BenchGen agent, medium effort, 100 steps, offline.v0.1 · updated 3 Oct 2026 · changelogRequest the full task set →
How to read this table
- pass@1
- Share of attempts where every hidden fail-to-pass test passes and no pass-to-pass test regresses, averaged per task.
- Hard
- pass@1 on the 55 tasks the calibration models usually missed; the subset that separates models.
- $ / solved task
- Average model API cost per attempt divided by pass@1. What a success costs, not an attempt.
- Interval
- 95% bootstrap over tasks. Shown as the thin line under each score.
- Rank
- Scale's rule: 1 + the number of models whose lower bound is above this model's upper bound. Overlapping intervals share a rank.
- Protocol
- BenchGen agent, medium reasoning effort, 100-step limit, no network. The patch present at the limit is graded as-is, as in DeepSWE. Every attempt against the shipped version of a task counts; attempts against earlier versions are discarded.
- Scheduled
- Claude Opus 5, GPT-5.6 Sol and other frontier models, not yet run.
What the tasks are
Real race conditions from 106 Python projects: lost updates, deadlocks, leaked cancellation. 74 combine several upstream fixes.
Browse the 95 public tasks →How it is graded
Sealed repository at one commit, no network, hidden tests run in a clean verifier, binary reward. Every task verified three times before release.
Methodology →Run it
git clone https://github.com/versocr/swe-race uv tool install datacurve-pier pier run -p swe-race/tasks --agent mini-swe-agent --model <provider/model>