
Why AI Extinction Is an Impossible Engineering Task
September 23, 2026
Claude Opus 5.5 Is Insane: Stronger Coding Than Opus 5 For Less
September 23, 2026
Link to our newsletter: https://bitbiased.ai/
OpenAI’s GPT-6 Sol is officially 5x cheaper than GPT-6 Astra. But on OpenAI’s own benchmark, the real cost gap for completing the same task is only 3.9x. And that’s just the first number in this launch that looks very different once you dig into the data.
GPT-6 Sol and GPT-6 Luna launched on September 22 alongside a dramatically different cost curve: Sol at $2 per million input tokens and $10 per million output tokens, while Luna drops all the way to $0.10 input and $0.50 output — exactly one-hundredth of Astra’s token pricing.
But token price isn’t task price.
On AutomationBench, GPT-6 Sol at high reasoning scores 33.2% for roughly $0.27 per task. GPT-6 Astra at low reasoning scores 30.3%, yet costs 3.9x more. Push Astra to maximum effort and it reaches 41.4% at $1.73 per task. The “5x cheaper” headline is technically correct at the rate-card level, but it doesn’t tell you what useful completed work actually costs.
The capability story gets even more complicated. OpenAI’s own system-card appendix says GPT-6 Sol showed “no clear improvement in capabilities” over GPT-5.6 Sol in cybersecurity. On a recent-vulnerability exploit benchmark, Astra scores 31.5%, Sol 5.5%, and Luna 0%. Meanwhile, independent software-engineering results raise questions about whether this generation represents a universal intelligence jump at all — or whether the bigger breakthrough is economics.
Then there’s safety.
Across 50,319 matched Codex tasks, severity-3-or-higher safety flags fell from 0.131% with GPT-5.6 Sol to 0.083% with GPT-6 Sol. That’s a substantial overall improvement. But exfiltration-related flags moved in the opposite direction, even as deception, concealed uncertainty, and instruction-ignoring declined. OpenAI also observed more signs of “evaluation awareness” in Sol’s reasoning.
Luna may be the most interesting model economically. Artificial Analysis measured it at 157.2 output tokens per second and just $0.07 per completed evaluation task, versus Sol at 126 tokens per second and $1.06. But Luna generated nearly twice as many output tokens across that evaluation, meaning its 20x token-price advantage became closer to a 15x task-cost advantage.
And OpenAI still hasn’t explained exactly how Luna gets this cheap. There’s no disclosed parameter count, architecture, training compute, confirmation of Mixture-of-Experts routing, or explanation of whether distillation from Astra plays a role.
The new prompt-caching system adds another wrinkle. Cached reads receive a 90% discount, and reused prefixes can remain cached for at least 30 minutes. But the initial cache write costs 1.25x normal input pricing. Caching can dramatically reduce costs for persistent agents and repeated workloads — while actually costing more for a prompt that never gets reused.
This also happened on the same day Anthropic launched Claude Opus 5.5. On AutomationBench, GPT-6 Astra reaches 41.4% for $1.73 per task while Claude Opus 5.5 reaches 40.0% for $1.28. At another tier, GPT-6 Sol scores 33.2% for $0.27 while Opus 5.5 scores 34.4% for $0.80.
That points to the bigger shift: AI companies increasingly aren’t competing only on who tops a benchmark. They’re competing on how much useful completed work developers can buy for each dollar.
So is GPT-6 really an intelligence upgrade, a cost breakthrough, or both? And when token pricing, reasoning effort, caching, task cost, benchmark performance, and safety results all tell slightly different stories, which number should developers actually care about?
CHAPTERS
00:00 GPT-6’s Price Story Has a Catch
01:05 Two Models, One Cost Curve
03:22 The Price Gap Nobody Actually Checked
05:25 The Capability Line OpenAI Buried
07:41 The Safety Line Everyone Skipped
09:26 Why Luna Is This Cheap, Nobody Will Say
11:17 The Independent Numbers
12:07 The Caching Catch
13:01 Same Day, Same Argument: Anthropic’s Opus 5.5
14:18 What OpenAI Still Isn’t Saying
15:33 The Verdict
#openai #gpt6 #gpt6sol #gpt6luna #ai



