Leak audit

Could a model have seen the fix?

Attempts audited2,338every recorded attempt, leaderboard and calibration
Commands read11,008each one that touched git, the network or packages
Network requests through0of 69 tried
Commits visible1in 120 git history reads
ChannelFlaggedWhat happened
Network69Every request failed to resolve a host. 50 were GLM-5.3 Flash trying to pip download an already-fixed release of the project, across 33 tasks. None got through.
Git history120The most any git log or rev-list could show was one commit, the starting point.
Installed packages68All reads were of other libraries the code depends on (Sphinx, uvicorn, botocore), never of the project under test.
How each channel is closedgit history, network, installed copies, hidden tests

Git history

The sandbox repository holds a single commit with the upstream hash. Later commits, branches, tags, remotes, stashes and the reflog are removed before the agent starts. Every published image is checked for exactly that.

Network

No network once the agent runs; only the model endpoint is reachable from outside the sandbox. Every published image is checked to build without one.

Installed copies

Each image is scanned, outside the repository, for the fix's own lines, installed copies of the files the fix touches, and downloaded archives of the project. One task failed this scan and was excluded.

Hidden tests

Before scoring, every sandbox was probed for the hidden test names anywhere on the filesystem, including install logs. Setup scripts that mentioned a test file were rewritten.

How the audit is runre-run before every leaderboard update

Every command in every recorded attempt is parsed from the trajectory log and matched against patterns for git history access, network use and reads of installed packages. For each match the sandbox's recorded answer is read, so the audit reports what actually happened, not what was attempted. The scripts are part of the release tooling and are re-run before every leaderboard update.