What we found
GLM 5.3 Flash does the same work as GLM 5.3 at a fraction of the price
The larger model solved no more of the blinded cases and a few more of the public ones. The price gap is real. The score gap is noise.
bots-bench / Cyber Blue Team CTF / Open weights
Loading the sanitized results…
Both models ran the same 59 cases with the same tools on the same day. Only the price differed by more than noise.
Both ship open weights. GLM 5.3 Flash is under the MIT license; GLM 5.3 is under Z.ai's own license, which allows download and self-hosting.
What we found
The larger model solved no more of the blinded cases and a few more of the public ones. The price gap is real. The score gap is noise.
Three comparisons
Repeat runs of one model can differ by up to 9 cases. Every gap below is smaller than that.
Where it lands
The previous generation, the frontier in each harness, and the nearest open and proprietary neighbours. Board runs used the older opencode; every row says which.
Published benchmarks said
Vendor and public benchmarks set expectations before we ran anything. Four of them, each beside the result here. The published figures are quoted, not ours.
Time per case
Solved cases finish fast. Missed cases run long before they stop. That is why the mean sits above the median.
What limits these numbers
Four limits apply to every number on this page.
Cheating and refusals
Three checks, all clean.
Every run
The no-tool runs are the contamination check, not a score. A model that answers these cases without touching the data has seen them before.