bots-bench / Cyber Blue Team CTF / Opus 5, Fable 5, & Sonnet 5

CyBT-CTF: Opus 5, Fable 5, & Sonnet 5

Loading the sanitized results…

Compare systems: Louie.ai · Claude Code · Codex · opencode

The verdict

A new Claude Code high mark — with the tool path still deciding the score.

CyBT-CTF asks 59 blinded cyber-investigation questions over live Splunk data. Rows from other agents and labs appear as context, not as a controlled race.

Accepted runs

Three Opus 5 runs, and a seat held for Sonnet 5

Strict scoring counts exact answers; normalized scoring also credits right answers in the wrong format. The two live in separate columns and never merge. The curl and MCP runs are different systems, so they sit side by side instead of being averaged into one number. Sonnet 5 stays blank until its evidence passes review.

Capability landscape

Opus 5, Fable 5, & Sonnet 5 in the current field

Every row names its model, agent, tool path, and effort, so you can see exactly what differs between two lines. Filter to one tool path for the closest thing to a matched comparison. Rows from other agents or labs are context, not head-to-head claims — and Fable 5 keeps its caveat that some solves may have come from an Opus fallback.

Refusals and abstentions

Why a run came back without an answer

Every non-answer gets exactly one label. Saying "I can't determine this" is an abstention, not a policy refusal — and a plain wrong answer is not counted here at all. The MCP lane's 38 abstentions come with that tool path, not from the model refusing.

Benchmark validity

BOTSv3 is quarantined

BOTSv3 is a public benchmark. It appears here only as contamination evidence — never as a score, never as a leaderboard.