Subset 17-ID control · with tools
- Opus 4.6 high: 14/17
- Opus 4.7 high: 5/17
- Paired delta: +9 for 4.6 on same IDs, same envelope, post-window.
Matched envelope: MCP, parallel=2, timeout=360s, no session persistence, high effort.
bots-bench / AI SOC evals / Splunk BOTSv3 + CyBT-CTF
Loading benchmark corpus and run snapshots…
Primary Benchmark View
CyBT-CTF is a set of cyber investigations the models have never seen, so it is the main comparison here. BOTSv3 is an older public test the models may have trained on, so we keep it as a cross-check you can switch to — not the headline.
BOTSv3 Frontier Leaderboard
Everything from here down covers the public BOTSv3 test. Each entry pairs an AI agent (Claude Code, OpenAI Codex, opencode, Louie) with the model it drives. For the main, blinded comparison, use the CyBT-CTF table above.
BOTSv3 Follow-Up
Each shape shows one setup per model family — the best-scoring run of its newest version.
Fable Follow-Up
Fable appears here only when the public data includes its refusal and quality numbers.
BOTSv3 model progress over time
One dot per model, keeping its strongest setting. Dots connect in release order within each family, so you can watch each generation move.
BOTSv3 Efficiency
BOTSv3 only. Switch between cost vs. speed, cost vs. solve rate, and cost vs. combined score. Bubble size shows how many tasks each setup solved on the first try, which keeps cost free for the axes.
Benchmark Program
Systems Under Test
BOTSv3 Benchmark Atlas
This is the public BOTSv3 collection, not CyBT-CTF. These are real investigations, not toy prompts: more than 100 overlapping log and alert sources, incidents drawn from the BOTSv3 case files, and questions that run from easy to hard.
Chart Wall
Each dot is one question. Farther left, fewer setups solve it. Higher up, it costs more time or more tries. Bigger dots took more repeated attempts. Color marks the track.
Claude 4.6 Controversy
Anthropic published an April 23 postmortem describing three product- and harness-level problems that hit recent Claude runs. We reran the same 17 previously-contaminated questions under matched conditions, to separate what the model did from quirks of timing.
Matched envelope: MCP, parallel=2, timeout=360s, no session persistence, high effort.
q2_sq201_id_cl solves for 4.7 — general-knowledge-answerable.Evidence: Anthropic's postmortem. Our matched-rerun findings and short report stay in the private benchmark repository, because they quote the questions and answers.
BOTSv3 Benchmark Hygiene
An answer-only control tests whether a model can reproduce BOTSv3 answers without using Splunk. Item review separates confirmed matches from screening signals; missing reviews remain unknown.
BOTSv3 Harness Versions
For each model, we compare its earliest and latest BOTSv3 runs — a rough way to see how the test setup changed over time.
BOTSv3 Prompt Effects
Looking for clean BOTSv3 cross-validation and OODA / OSCAR-style planning-loop pairs...
BOTSv3 Latest Runs
Looking for local run artifacts…
BOTSv3 Retry Lab · handle with care
Some setups let a model try again after a wrong answer. Extra tries can nudge the score up a little — but they burn far more time and money for that small gain. This section is flagged in red because those scores aren't a fair comparison to single-attempt runs.
BOTSv3 Config Microscope
The map and the microscope share one selection. Pick a setup here, or click a dot in the solve-vs-time map, to see how it did question by question instead of just the totals.
Config Map
Plain setup names, not raw run IDs. Choose the score and time measures you care about, then click a dot to load it in the microscope below.
BOTSv3 Supporting Boards
Additional ways to slice the same public BOTSv3 snapshot.
BOTSv3 Benchmark Rules
BOTSv3 Experiment Surface
What we change from run to run on BOTSv3. CyBT-CTF is tracked separately, on the comparison and deep-dive pages.