bots-bench / BOTSv3 benchmark hygiene

Could the model already know the BOTSv3 answers?

A model that reproduces BOTSv3 answers without using Splunk may have seen them before. Item review separates confirmed matches from screening signals; missing reviews remain unknown.

Only aggregate audit results are public. Audit prompts, answers, traces, and run IDs stay private.

Where this stands

Contamination review

By model

What we found, model by model

How much each model answered from memory alone, and whether anyone has reviewed those answers.

All controls

Every model line in one sortable table

"Not reviewed" and "not tested" mean the check has not run. That is never the same as a clean result.