bots-bench / Cyber Blue Team CTF / Open weights
CyBT-CTF: GLM 5.2
GLM 5.2 is the strongest open-weight model on CyBT-CTF. It solves 35 of 59 tasks — tying the unlocked Fable's full run, three ahead of Claude Code / Opus 4.8 (32), and one ahead of the best stock-Codex GPT-5.6 run (Luna high, 34), though Luna solves each task for half of GLM's cost. The bargain of the tier is Qwen3.7 Plus, a proprietary Alibaba model: two solves behind at 33, for a quarter of GLM's cost per solve and a quarter of its time.
What makes GLM stand out is how closely its answers track GPT-5.5's: MCC 0.721 and Cohen's kappa 0.719 on the blinded test, under our corrected July 20 rubric. High enough to warrant a look at possible distillation or copying — but not proof on its own, since the two never gave the same wrong answer word-for-word (0 matches).