bots-bench / Cyber Blue Team CTF / Open weights

GLM 5.3 Flash: GLM 5.3's score for a fraction of the price

Loading the sanitized results…

Both models ran the same 59 cases with the same tools on the same day. Only the price differed by more than noise.

Both ship open weights. GLM 5.3 Flash is under the MIT license; GLM 5.3 is under Z.ai's own license, which allows download and self-hosting.

What we found

GLM 5.3 Flash does the same work as GLM 5.3 at a fraction of the price

The larger model solved no more of the blinded cases and a few more of the public ones. The price gap is real. The score gap is noise.

Compare systems: Louie.ai · Claude Code · Codex · opencode

Three comparisons

What the gaps mean

Repeat runs of one model can differ by up to 9 cases. Every gap below is smaller than that.

Where it lands

GLM 5.3 beside the runs a reader will compare it with

The previous generation, the frontier in each harness, and the nearest open and proprietary neighbours. Board runs used the older opencode; every row says which.

Published benchmarks said

What we expected, and what the blinded cases showed

Vendor and public benchmarks set expectations before we ran anything. Four of them, each beside the result here. The published figures are quoted, not ours.

Time per case

Failure costs more time than success

Solved cases finish fast. Missed cases run long before they stop. That is why the mean sits above the median.

What limits these numbers

Read this before quoting a score

Four limits apply to every number on this page.

Cheating and refusals

None found

Three checks, all clean.

Every run

With tools, and without

The no-tool runs are the contamination check, not a score. A model that answers these cases without touching the data has seen them before.