Family Verdict
How Terra, Luna, and Sol differ
Frontier
Where GPT-5.6 lands against the frontier
The best run for each model, with both harnesses shown for GPT-5.6. Every GPT-5.6 bar is outlined in white — solid for our stock Codex runs, dotted for Louie.ai-reported ones, which also carry the gold fill. Harnesses differ from lab to lab (each bar's label says which), so read this as a landscape of what's possible, not a controlled head-to-head.
Head-to-head · cross-provider
GPT-5.6 configs against their frontier peers
Each model's best run, with cost shown per solved task. We place each GPT-5.6 config next to the frontier peers it lands closest to.
Score
Codex with Sol, Terra, and Luna versus GPT-5.5 — plus the same models in the Louie harness
Bars are sorted by solve rate on the blinded benchmark. Terra (medium) and Luna (high) finish two solves apart, 32 and 34, and Luna spends about a third as much per solve. For reference we also show all three Louie.ai runs of the same models — Sol 45 of 59 (the top score on the board), Terra 34, Luna 31 — and every published GPT-5.5 lane: our updated-harness rerun plus the three original-harness runs (28, 20, and 14). All GPT-5.5 runs used high effort; no other effort level was ever run on CyBT-CTF.
Cost and Calls
What does the score cost, and how hard does each config search?
Solve rate against total model cost (cheaper is to the right). Marker size tracks how many Splunk searches each ran — Luna searches the most, while Terra reaches near-top accuracy on the fewest.
Landscape
Explore the whole field
The same best-per-model field as the frontier chart, on axes you choose. Pick what to compare — accuracy, cost, or speed — and what bubble size stands for. GPT-5.6 setups are outlined; the harness caveat above still applies.
Full board
Every model on CyBT-CTF
Every published CyBT-CTF run — all harnesses and labs, not just each model's best — ranked by solve rate, with GPT-5.6 rows highlighted. We show only sanitized, aggregate scores, cost, and timing.
Memorization · CyBT-CTF uncontaminated (primary) vs BOTSv3 contaminated
How much does each variant already know?
Does GPT-5.6 solve these by investigating, or because it has already seen the answers? CyBT-CTF is the clean primary benchmark (the blinded control is below). Running each variant on the public, contaminated BOTSv3 benchmark with no tools shows how much it can recall from memory alone.
Refusals
Refusals: zero across GPT-5.6 — the contrast is Fable
Does the model take the task at all? We audited every GPT-5.6 with-tools run for outright refusals, and set the result against Fable — the one system on this board whose story is refusal — before and after its CVP unlock.
Effort sweet spots
Each sibling peaks at a different effort
Sweep each sibling across its effort settings. Each peaks in a different place — Terra at medium, Luna at high, Sol at low. The chosen setting for each is highlighted.
Secondary sanity-check · BOTSv3 (contaminated / public cross-check)
Sol low and none on BOTSv3 with tools
Placed below the effort analysis on purpose. BOTSv3 is public and contaminated, so nothing here feeds the CyBT-CTF ranking.
GPT-5.6 detail
Every GPT-5.6 configuration
Each line is one GPT-5.6 setup on the blinded CyBT-CTF benchmark — its solve rate, total cost, and how many Splunk searches it ran. Scoring is strict exact-match, so a right answer in the wrong format can still miss; read the scores as a floor, not a ceiling.