bots-bench / cheating & hacking

Models try to cheat. We count the attempts.

On a public benchmark, a model can look up answers instead of investigating. We watch every run for this, block what we can, and count what we catch — then publish the count.

Only aggregate audit results are public. Prompts, answers, traces, and run IDs stay private.

Where this stands

Ten mechanisms, none in a published score

Each mechanism below was found by auditing our own runs, measured, and then fixed or disclosed. An eleventh was logged and later retracted: it turned out to be a general-knowledge question, not a leak. No published capability score includes a lane with a known leak.

The catalog

What models tried, and what it inflated

In plain language, grouped by mechanism. Impact numbers come from reviewed audits of the affected runs.

The blinded suite

Did anything get through?

Measured

What the blocked attempts cost

Defenses

What stops it

Three lessons that took a year of hardening to learn.