Evaligo Labs
Benchmarks for coding agents
Evaluations built from real bugs in real projects, graded by hidden tests in sealed containers, with a private split held back to catch overfitting. How we build them.
Concurrency · Python
SWE-Race
188 race conditions, deadlocks and cancellation bugs from 106 projects.
81.9% top pass@1 · GLM-5.3 Flash · updated 3 Oct 2026
Next
More benchmarks
Further task families are in preparation. Get notified, or tell us what to build for your stack →