Submit a model
How a model gets on the board
Who runs itWe doone protocol for every model; no self-reported scores
A full run188 × 4tasks × attempts; public and private together
Cost to you$0we report the API cost back to you
Protocol
| Harness | BenchGen agent: shell tool, file edits, submit |
| Step limit | 100 tool calls; the patch at the limit is graded as-is |
| Command timeout | 600 s |
| Reasoning effort | medium, where available |
| Network | none |
| Attempts | 4 per task; pass@1 is the mean per task |
| Grading | hidden tests in a separate verifier container; binary reward |
Steps
- Send the request form with the model identifier, provider and the inference settings you consider standard.
- For a private endpoint, include access for the run. Keys are used for the evaluation only and deleted afterwards.
- We run 188 tasks × 4 attempts, public and private, under the protocol on the left.
- You get the per-attempt results and the cost; the board updates and subscribers are emailed.
Run it yourselfpublic split, any Harbor runner
The public split is on GitHub and Hugging Face, as a Harbor task set that runs under Pier or any Harbor-compatible runner.
git clone https://github.com/versocr/swe-race uv tool install datacurve-pier pier run -p swe-race/tasks --agent mini-swe-agent --model <provider/model>
Your own numbers on the public split are welcome, marked as self-reported; they are listed separately from the board.
CiteBibTeX
@misc{sweracev01,
title = {SWE-Race: a benchmark of real concurrency bugs for coding agents},
author = {Evaligo Labs},
year = {2026},
url = {https://labs.evaligo.com/swe-race}
}