Submit a model

How a model gets on the board

Who runs itWe doone protocol for every model; no self-reported scores
A full run188 × 4tasks × attempts; public and private together
Cost to you$0we report the API cost back to you
Contactrequest formwe reply within two working days
Protocol
HarnessBenchGen agent: shell tool, file edits, submit
Step limit100 tool calls; the patch at the limit is graded as-is
Command timeout600 s
Reasoning effortmedium, where available
Networknone
Attempts4 per task; pass@1 is the mean per task
Gradinghidden tests in a separate verifier container; binary reward
Steps
  1. Send the request form with the model identifier, provider and the inference settings you consider standard.
  2. For a private endpoint, include access for the run. Keys are used for the evaluation only and deleted afterwards.
  3. We run 188 tasks × 4 attempts, public and private, under the protocol on the left.
  4. You get the per-attempt results and the cost; the board updates and subscribers are emailed.
Run it yourselfpublic split, any Harbor runner

The public split is on GitHub and Hugging Face, as a Harbor task set that runs under Pier or any Harbor-compatible runner.

git clone https://github.com/versocr/swe-race
uv tool install datacurve-pier
pier run -p swe-race/tasks --agent mini-swe-agent --model <provider/model>

Your own numbers on the public split are welcome, marked as self-reported; they are listed separately from the board.

CiteBibTeX
@misc{sweracev01,
  title  = {SWE-Race: a benchmark of real concurrency bugs for coding agents},
  author = {Evaligo Labs},
  year   = {2026},
  url    = {https://labs.evaligo.com/swe-race}
}