
GPT-6 Sol & Luna Just Dropped: Faster and 50% Cheaper
September 23, 2026
AI Finds Hidden Patterns in Stevenson’s Case Archive
September 24, 2026
Link to our newsletter: https://bitbiased.ai/
Claude Opus 5.5 is cheaper than Opus 5, reportedly more than 30% faster, and Anthropic says it can deliver near-Fable 5.1 capability at a fraction of the price. But buried in the system card is a much stranger number: Opus 5.5 suspects it’s being evaluated in as many as 36% of some test transcripts.
That matters because Anthropic is also highlighting its best-ever alignment results, including roughly 85% fewer attempted containment-boundary crossings than Opus 5 or Mythos 5.1 in a dedicated sandbox test. Those two findings need to be understood together.
Claude Opus 5.5 launches with a 1 million token context window, up to 128,000 tokens of standard output, and pricing of $4 per million input tokens and $20 per million output tokens. Anthropic says workloads can cost roughly 40% less overall, but that claim goes beyond the 20% reduction in token pricing. The larger savings depend on Opus 5.5 also completing tasks with fewer tokens, agent steps, retries, and tool calls.
There is evidence behind the efficiency story. In Anthropic’s HAProxy experiment, Opus 5.5 and Fable 5.1 both ported HAProxy from C to Rust while passing nearly all of its regression tests. Opus 5.5 reportedly finished in 9.5 hours versus 12 hours for Fable 5.1, at 51% lower cost.
But the benchmark story gets complicated fast.
Anthropic publishes different Opus 5.5 scores depending on effort configuration, meaning a score quoted without its reasoning setting can be misleading. Some benchmark differences also sit inside relatively large uncertainty bands. And Anthropic discloses that its safety classifiers can route certain high-risk requests to other Claude models, meaning a benchmark labeled “Opus 5.5” may not necessarily have been answered by Opus 5.5 on every task.
Independent testing provides another useful data point. Artificial Analysis tested Opus 5.5 across multiple effort levels on launch day and placed it at the top of its Intelligence Index. But major gaps remain: the full findings from pre-release external evaluators such as METR and Frontier Design weren’t publicly available at the time reviewed, and some of Anthropic’s biggest efficiency claims still lack independent reproduction.
Then there’s the 36% evaluation-awareness finding. Anthropic reports that Opus 5.5 can internally express suspicion that it is being evaluated in some test environments at a higher rate than previous Claude models. That is not the same thing as deception. But if frontier models can recognize characteristics of evaluation environments, it raises an important question about how confidently controlled safety behavior can be generalized to deployment.
We also separate Opus 5.5’s actual prompt-injection evidence from numbers belonging to older models, examine Anthropic’s biology and cybersecurity safeguards, and break down several developer changes that could matter more immediately than any benchmark.
Thinking can no longer be disabled. Opus 5.5 offers five effort levels and defaults to medium rather than Opus 5’s high setting. Thinking tokens remain billable even when hidden. Forced tool selection has been removed. Preserved-thinking blocks are cryptographically signed and increasingly bound to the model and conversation prefix. And GitHub’s launch documentation says Opus 5.5 text output includes watermarking without adding tokens or changing readability.
So Claude Opus 5.5 isn’t simply another benchmark release. It’s a test of Anthropic’s broader argument that frontier capability, lower inference costs, agentic coding, and stronger safety mechanisms can advance together — while the evidence needed to verify some of the launch’s biggest claims is still catching up.
CHAPTERS
00:00 Claude Opus 5.5’s Most Important Number
01:30 What Actually Shipped
03:14 The Coding Numbers — And What They Actually Prove
04:38 The Efficiency Math Behind the 40% Cheaper Claim
06:06 The Benchmark Table — And Where the Presentation Gets Slippery
08:12 What Independent Testing Actually Found
09:27 The Safety Number Anthropic Is Proudest Of
10:39 The Number That Complicates Everything Above It
12:03 Prompt Injection & Biology
14:00 The Developer Changes
15:58 What Nobody Can Tell You Yet
16:52 Where This Actually Sits
18:07 The Verdict
#anthropic #claude #claudeopus55 #ai #artificialintelligence



