bots-bench / Cyber Blue Team CTF

Fable, unlocked: it edges Opus at cyber defense — when it agrees to answer

Loading Fable aggregate view...

Compare systems: Louie.ai · Claude Code · Codex · opencode

Reader Questions

What this page is trying to separate

Leaderboard

Score, solve time, and cost on the selected test

This chart ranks both Fable versions we measured — locked and unlocked — plus "Mythos", our projection of a Fable that never refuses, alongside Opus, Codex, and top open-weight systems, one test at a time. CyBT-CTF is the main, blinded test; BOTSv3 is a public cross-check.

Solve Rate

Solve rates on the selected test

Solid bars are measured. Dotted "Mythos" bars estimate how a never-refusing Fable would do, based on same-task evidence from the unlocked Fable versus Opus. Bars are sorted so better results sit farther right.

Tradeoff Map

Interactive score, time, and cost scatter

Use the toggles to compare accuracy, solve time, median time, and cost. Where lower is better, we flip the axis so the better points still move right or up.

Who Answered

How often Fable answered for itself versus handing off to Opus

Only Fable ever hands a task to another model, so this chart tracks Fable alone. Each bar is one measured run — both the locked and the unlocked version — showing how often that Fable answered on its own, handed the task to Opus, or returned no answer.

Coverage

Which attack tactics each system handles (ATT&CK skill map)

Proprietary Systems

Unlocked Fable vs Opus and Codex

Open-Weight Systems

Unlocked Fable vs open-weight models

The Refusals

Which ordinary cyber tasks the unlocked Fable still declined

This is the main caveat. Even unlocked, Fable refused a handful of ordinary cyber-defense tasks, and our harness handed those to Opus to finish; timeouts and unfinished attempts are counted separately. The locked Fable's records stay below as history.

Answer Path

Answered by Fable vs finished by Opus vs incomplete

Reasons

Why attempts were passed to Opus

Benchmark Validity

CyBT-CTF versus public BOTSv3

CyBT-CTF is the main test because the models have never seen it. BOTSv3 is public, so it is useful only as a contamination check: if a model jumps on BOTSv3 but not on CyBT-CTF, read that as a warning sign, not a real gain in skill.