bots-bench / Cyber Blue Team CTF / Open weights
CyBT-CTF: GLM 5.2
GLM 5.2 solves 33 of 59 tasks, the strongest open weight on this board, one behind the best stock-Codex GPT-5.6 run (Luna high, 34), and two behind the unlocked Fable's full run (35). Its successor GLM 5.3 Flash has since gone ahead of it. The bargain of the tier is Qwen3.7 Plus, a proprietary Alibaba model: the same 33, for a quarter of GLM's cost per solve and a quarter of its time.
Update, September 2026: Z.ai has released GLM 5.3 and GLM 5.3 Flash. We measured both. Flash does the same work as GLM 5.3 at a fraction of the price, and in Louie.ai it reports 38 of 59 to GLM 5.2's 37 there. They ran on a newer opencode, so they have their own page rather than a row here.
What makes GLM stand out is how closely its answers track GPT-5.5's: MCC 0.721 and Cohen's kappa 0.719 on the blinded test, under our corrected July 20 rubric. High enough to warrant a look at possible distillation or copying — but not proof on its own, since the two never gave the same wrong answer word-for-word (0 matches).