bots-bench / Cyber Blue Team CTF / Open weights

CyBT-CTF: GLM 5.2

GLM 5.2 is the strongest open-weight model on CyBT-CTF. It solves 35 of 59 tasks — tying the unlocked Fable's full run, three ahead of Claude Code / Opus 4.8 (32), and one ahead of the best stock-Codex GPT-5.6 run (Luna high, 34), though Luna solves each task for half of GLM's cost. The bargain of the tier is Qwen3.7 Plus, a proprietary Alibaba model: two solves behind at 33, for a quarter of GLM's cost per solve and a quarter of its time.

What makes GLM stand out is how closely its answers track GPT-5.5's: MCC 0.721 and Cohen's kappa 0.719 on the blinded test, under our corrected July 20 rubric. High enough to warrant a look at possible distillation or copying — but not proof on its own, since the two never gave the same wrong answer word-for-word (0 matches).

Frontier comparison

GLM 5.2 leads open weights at 35/59 — but Qwen3.7 Plus gets 33 for a quarter of the cost

On the blinded test, GLM 5.2 solves 35 of 59 and Qwen3.7 Plus 33 — each run through opencode on Fireworks, but Qwen gets there for a quarter of the cost. GLM stays the focus of the agreement audit because its answers track GPT-5.5's so closely.

Compare systems: Louie.ai · Claude Code · Codex · opencode

Executive Read

Where GLM fits

Both tabs show the same models on different tests. CyBT-CTF is blinded, so it carries the verdict; public BOTSv3 can reward memory.

Score

GLM versus Opus, Sonnet, Codex, and open-weight models

Bars are sorted by solve rate. GLM is highlighted. When scores tie, the table says so and orders those entries by solve time, then cost. Estimated and raw private results are left out.

Cost and Time

How much does the score cost?

The scatter shows solve rate versus total time spent on solved tasks. Marker size follows total model cost when available.

Agreement Audit

Does GLM look unusually close to frontier closed models?

This is a short extract from the full agreement audit. When two models from different companies agree this often — especially on wrong answers — it can be a sign to check for distillation or copying.

CyBT-CTF same-benchmark triangle

GLM 5.2 vs GPT-5.5 is the standout edge

MCC 0.763 Kappa 0.763 MCC 0.721 Kappa 0.719 MCC 0.630 Kappa 0.626 Opus 4.8 opencode GLM 5.2 Fireworks GPT-5.5 Codex

Strongest review signal

Z.ai / GLM 5.2 vs OpenAI / GPT-5.5

0.719 Cohen's kappa
MCC / phi
0.721
Solved overlap
80.0%
Observed / expected
1.95x
Same wrong calls
10 canonical / 0 exact text

A kappa above 0.70 is unusually high for systems built by different companies on the same blinded tasks. Treat it as a serious prompt to review for distillation or copying — not a public accusation, and not proof.

Open the full agreement audit

Leaderboard

What's behind the charts

We show only sanitized, aggregate results. Tied ranks share the same solve count; within a tie, we order by solve time, then total cost. No task text, raw traces, prompts, answers, SPL queries, or private benchmark markers appear here.