bots-bench / agreement audit

Model Agreement Audit

Agreement cuts both ways. Within a model family, it shows which behaviors carry over from one generation — or one sibling — to the next. Across companies, unusually close behavior can flag possible distillation, copying, or benchmark contamination. Treat this page as a screening view, not proof on its own.

Updated July 20, 2026: we corrected the scoring rubric and rescored every frozen answer. Correlations came down from about 0.80 to about 0.72 — still objectively high for models built by different companies, and still worth a look. GPT-5.6 (Terra, Luna, Sol) joins the table for the first time.

Compare systems: Louie.ai · Claude Code · Codex · opencode

Executive Read

What agreement can and cannot say

Use this as a first-pass screen. The good case is continuity: sibling and next-generation models should keep some useful behavior, and this shows where that holds or breaks. The worrying case is imitation: when two models from different companies look too alike — especially on wrong answers — the pair deserves a private follow-up.

MCC / phi

Matthews correlation coefficient is a -1 to +1 agreement score over what two systems solved and missed. Near 1 means very similar outcomes, near 0 means little relationship, and negative means opposite patterns. We use it as the headline affinity score because it handles uneven solve rates better than raw overlap.

Cohen’s kappa

Cohen’s kappa is another chance-adjusted agreement score. It is useful as a second read: if MCC and kappa both look high, the agreement signal is stronger; if they diverge, the pair needs closer review.

Featured Triangles

Cross-company screens and same-company controls

This shows one same-benchmark triangle at a time. CyBT-CTF is the blinded view; BOTSv3 is the public, contamination-prone cross-check — and the only place here with a full GPT-5.3-codex run. Controls pair models from one company (Opus 5 with Sonnet 5, Terra with Luna) so you can see what normal same-company agreement looks like before judging a cross-company match. Line width shows how strong the match is (MCC/phi), and color marks how much follow-up the link deserves.

MCC / phi edges

Same-benchmark affinity triangles

Line width shows the match strength (MCC/phi). Orange flags an unusually high match across companies, yellow an elevated one, and blue the expected similarity inside a single company or family.

Same-Benchmark Pairs

Highest-affinity pairs

The main table never ranks BOTSv3 and CyBT-CTF against each other. CyBT-CTF is the blinded benchmark; BOTSv3 is public and contamination-prone. GPT-5.3-codex shows up only under BOTSv3, because there is no full, safe CyBT-CTF run for it. On CyBT-CTF, we now include the latest GPT-5.6 (Terra, Luna, Sol) and Fable runs against the references we can source directly — Opus 4.8, GPT-5.5, and GLM 5.2; Qwen3.7 Plus and Sonnet pairings still await a fuller rerun. Use "Compare against" to pick any model and rank the rest by how closely they agree with it.

MCC / phi

Top same-benchmark affinity

Higher MCC means the two systems solved and missed more of the same tasks. Similarity within a family is expected; similarity across companies is a review signal when it beats the controls or clusters on wrong answers.

Overlap Map

Affinity versus solved overlap

Same-company and same-family points show normal continuity. Cross-company points help you spot behavior that may deserve a private review for distillation, contamination, or a leaked evaluation. Hover or click any point for the exact pair; the strongest signals are labeled.

Report-Level Signal

Closest match inside and outside each company

Every system shows two bars: its closest match inside its own company, and its closest match outside it. Bars group by company. Hover a bar to see which system produced that match. When the outside bar beats the inside one, that pair is worth a closer look — not a conclusion.

Provider Groups

Provider and family summaries

These summaries cover same-benchmark pairs only. Use the maximum to find where to look, and the mean so you don't over-read any single pair.