bots-bench / Cyber Blue Team CTF / Open weights

CyBT-CTF: GLM 5.2

GLM 5.2 solves 33 of 59 tasks, the strongest open weight on this board, one behind the best stock-Codex GPT-5.6 run (Luna high, 34), and two behind the unlocked Fable's full run (35). Its successor GLM 5.3 Flash has since gone ahead of it. The bargain of the tier is Qwen3.7 Plus, a proprietary Alibaba model: the same 33, for a quarter of GLM's cost per solve and a quarter of its time.

Update, September 2026: Z.ai has released GLM 5.3 and GLM 5.3 Flash. We measured both. Flash does the same work as GLM 5.3 at a fraction of the price, and in Louie.ai it reports 38 of 59 to GLM 5.2's 37 there. They ran on a newer opencode, so they have their own page rather than a row here.

What makes GLM stand out is how closely its answers track GPT-5.5's: MCC 0.721 and Cohen's kappa 0.719 on the blinded test, under our corrected July 20 rubric. High enough to warrant a look at possible distillation or copying — but not proof on its own, since the two never gave the same wrong answer word-for-word (0 matches).

Frontier comparison

GLM 5.2 leads the open weights on this board at 33/59 — and proprietary Qwen3.7 Plus ties it for a quarter of the cost

On the blinded test, GLM 5.2, Kimi K3 and Qwen3.7 Plus each solve 33 of 59 — all run through opencode on Fireworks, but Qwen gets there for a quarter of the cost. GLM stays the focus of the agreement audit because its answers track GPT-5.5's so closely. GLM 5.3 and GLM 5.3 Flash ran later on a newer opencode. They have their own page.

Compare systems: Louie.ai · Claude Code · Codex · opencode

Executive Read

Where GLM fits

Both tabs show the same models on different tests. CyBT-CTF is blinded, so it carries the verdict; public BOTSv3 can reward memory.

Score

GLM versus Opus, Sonnet, Codex, and open-weight models

Bars are sorted by solve rate. GLM is highlighted. When scores tie, the table says so and orders those entries by solve time, then cost. Estimated and raw private results are left out.

Cost and Time

How much does the score cost?

The scatter shows solve rate versus total time spent on solved tasks. Marker size follows total model cost when available.

Agreement Audit

Does GLM look unusually close to frontier closed models?

This is a short extract from the full agreement audit. When two models from different companies agree this often — especially on wrong answers — it can be a sign to check for distillation or copying.

CyBT-CTF same-benchmark triangle

GLM 5.2 vs GPT-5.5 is the standout edge

MCC 0.763 Kappa 0.763 MCC 0.721 Kappa 0.719 MCC 0.630 Kappa 0.626 Opus 4.8 opencode GLM 5.2 Fireworks GPT-5.5 Codex

Strongest review signal

Z.ai / GLM 5.2 vs OpenAI / GPT-5.5

0.719 Cohen's kappa
MCC / phi
0.721
Solved overlap
80.0%
Observed / expected
1.95x
Same wrong calls
10 canonical / 0 exact text

A kappa above 0.70 is unusually high for systems built by different companies on the same blinded tasks. Treat it as a serious prompt to review for distillation or copying — not a public accusation, and not proof.

Open the full agreement audit

Self-hosted

GLM 5.2 on your own hardware: four H200s, 4-bit

Louie.ai ran the same 59 blinded cases on a self-hosted, 4-bit GLM 5.2 and reported the aggregates. The box keeps up on speed. It loses cases to the time limit, not to wrong answers.

Leaderboard

What's behind the charts

Tied ranks share the same solve count; within a tie, we order by solve time, then total cost.