bots-bench / Cyber Blue Team CTF / GPT-6 Astra & Claude Fable 5.1

Astra gains accuracy. Fable 5.1 stops refusing.

Loading the results…

Compare systems: Louie.ai · Claude Code · Codex · opencode

Leaderboard

Every system, the same cases, the same tools

Cost

Accuracy against cost per case

Effort

More thinking helps Astra, not Fable

Deadline

Both models now finish inside the time limit

Refusals

Neither model declined a single case

Integrity

No cheating, no refusals, no downgrades

The site-wide checks live on their own pages: Cheating · Refusals · Hygiene · About & methods.