bots-bench / AI SOC evals / Splunk BOTSv3 + CyBT-CTF

Benchmarking AI models and agents for SOC and IR investigations.

Loading benchmark corpus and run snapshots…

Made with love by Team Graphistry, makers of Louie.ai

Primary Benchmark View

CyBT-CTF first, BOTSv3 as the public cross-check

CyBT-CTF is a set of cyber investigations the models have never seen, so it is the main comparison here. BOTSv3 is an older public test the models may have trained on, so we keep it as a cross-check you can switch to — not the headline.

BOTSv3 Frontier Leaderboard

Public BOTSv3 agent + model leaderboard

Everything from here down covers the public BOTSv3 test. Each entry pairs an AI agent (Claude Code, OpenAI Codex, opencode, Louie) with the model it drives. For the main, blinded comparison, use the CyBT-CTF table above.

BOTSv3 Follow-Up

Model coverage by ATT&CK tactic

Each shape shows one setup per model family — the best-scoring run of its newest version.

BOTSv3 model progress over time

How each family scales by release date

One dot per model, keeping its strongest setting. Dots connect in release order within each family, so you can watch each generation move.

BOTSv3 Efficiency

Public BOTSv3 efficiency tradeoff explorer

BOTSv3 only. Switch between cost vs. speed, cost vs. solve rate, and cost vs. combined score. Bubble size shows how many tasks each setup solved on the first try, which keeps cost free for the axes.

Benchmark Program

What this benchmark is, and what it refuses to fake.

Systems Under Test

Agent harnesses and models on the board

BOTSv3 Benchmark Atlas

The public BOTSv3 corpus underneath these scoreboards

This is the public BOTSv3 collection, not CyBT-CTF. These are real investigations, not toy prompts: more than 100 overlapping log and alert sources, incidents drawn from the BOTSv3 case files, and questions that run from easy to hard.

Track coverage

ATT&CK coverage

Chart Wall

Question hardness

Each dot is one question. Farther left, fewer setups solve it. Higher up, it costs more time or more tries. Bigger dots took more repeated attempts. Color marks the track.

Claude 4.6 Controversy

Anthropic Apr-23 postmortem: did we catch a regression?

Anthropic published an April 23 postmortem describing three product- and harness-level problems that hit recent Claude runs. We reran the same 17 previously-contaminated questions under matched conditions, to separate what the model did from quirks of timing.

Subset 17-ID control · with tools

  • Opus 4.6 high: 14/17
  • Opus 4.7 high: 5/17
  • Paired delta: +9 for 4.6 on same IDs, same envelope, post-window.

Matched envelope: MCP, parallel=2, timeout=360s, no session persistence, high effort.

Subset 17-ID · no tools

  • Opus 4.6: 6/17 (high), 7/17 (default)
  • Opus 4.7: 1/17 (both effort lanes)
  • Only q2_sq201_id_cl solves for 4.7 — general-knowledge-answerable.

Interpretation

  • Window-confound alone does not explain the 4.6 > 4.7 gap.
  • Residual gap likely includes model-behavior and tool-use differences.
  • Host-level config changes around Apr 17 remain a credible, unquantified confound for non-isolated runs.

Evidence: Anthropic's postmortem. Our matched-rerun findings and short report stay in the private benchmark repository, because they quote the questions and answers.

BOTSv3 Benchmark Hygiene

Could the model already know the BOTSv3 answers?

An answer-only control tests whether a model can reproduce BOTSv3 answers without using Splunk. Item review separates confirmed matches from screening signals; missing reviews remain unknown.

BOTSv3 Harness Versions

Same model across Claude Code MCP run dates on BOTSv3

For each model, we compare its earliest and latest BOTSv3 runs — a rough way to see how the test setup changed over time.

BOTSv3 Prompt Effects

Cross-validation and planning loops on BOTSv3

Looking for clean BOTSv3 cross-validation and OODA / OSCAR-style planning-loop pairs...

BOTSv3 Latest Runs

Pass-rate snapshot cards for published BOTSv3 executions

Looking for local run artifacts…

BOTSv3 Retry Lab · handle with care

Does giving a model more tries actually help?

Some setups let a model try again after a wrong answer. Extra tries can nudge the score up a little — but they burn far more time and money for that small gain. This section is flagged in red because those scores aren't a fair comparison to single-attempt runs.

BOTSv3 Config Microscope

Drill into one BOTSv3 configuration, one question at a time

The map and the microscope share one selection. Pick a setup here, or click a dot in the solve-vs-time map, to see how it did question by question instead of just the totals.

Config Map

Solve rate versus time

Plain setup names, not raw run IDs. Choose the score and time measures you care about, then click a dot to load it in the microscope below.

BOTSv3 Supporting Boards

Secondary cuts of the same snapshot

Additional ways to slice the same public BOTSv3 snapshot.

BOTSv3 Benchmark Rules

How the BOTSv3 runs work

BOTSv3 Experiment Surface

What we are varying

What we change from run to run on BOTSv3. CyBT-CTF is tracked separately, on the comparison and deep-dive pages.