
Diamonds in the Rack: Room-Temp Quantum Computing – with Marius Grundmann of SAXON Q | Ep. 137
September 22, 2026
Gemini 4 Pro Just Leaked: The Results Are Insane
September 22, 2026
Link to our newsletter: https://bitbiased.ai/
Grok 4.7 just jumped from 31.9% to 46.0% on SWE-Marathon — a 44% relative improvement in just six weeks. But when independent lab Artificial Analysis tested xAI’s new model, its composite score moved from 44 to just 46… while generating more than twice as many output tokens.
That gap between xAI’s headline benchmarks and independent testing is where Grok 4.7 gets interesting.
xAI announced Grok 4.7 on September 21, 2026, just 40 days after Grok 4.6. According to xAI, this is a new, larger base model with longer supplemental training, curated reasoning data, advanced technical corpora, an improved training recipe, and heavier reinforcement learning on difficult multi-hour agentic tasks.
But most regular Grok users can’t actually use it yet. Grok 4.7 launched through the xAI API, Grok Build, Cursor, and Microsoft Office add-ins first, while access through Grok’s web and mobile apps and Grok on X is coming later.
The headline improvements are concentrated around coding and long-running agents. We break down DeepSWE, Terminal-Bench 4.0, FrontierSWE V2, SWE-Marathon, and CursorBench 4.0 individually — including where Grok 4.7 makes clean gains and where reasoning-effort differences make xAI’s comparisons less straightforward.
There are some impressive numbers. Grok 4.7 reaches 71.0% on DeepSWE and jumps to 46.0% on SWE-Marathon. But competitors including GPT-5.6 Sol, Claude Fable 5.1, and Opus 5 still lead several of these coding benchmarks.
The bigger story may be agents that can work for hours. xAI tested Grok 4.7 on software engineering tasks with time budgets reaching 20 hours and is increasingly positioning Grok for professional work beyond coding: legal documents, spreadsheets, presentations, CAD, electrical engineering, biological data, and other complex workflows.
Then comes the independent reality check.
Artificial Analysis scored Grok 4.7 high at 46 on its Intelligence Index, compared with 44 for Grok 4.6 high. But its evaluation produced roughly 200 million output tokens for 4.7 versus about 94 million for 4.6. Measured cost per evaluated task also increased from $1.86 to $2.73. Same list price per token does not necessarily mean the same cost per completed task.
Pricing itself remains $2 per million input tokens and $6 per million output tokens for standard requests at or below 200,000 tokens. Go beyond that threshold and the entire request moves to the higher long-context pricing tier. We compare that with GPT-6 Astra, GPT-5.6 Sol, Claude Fable 5.1, and Gemini 3.8 Flash — and look at whether Grok can still claim a meaningful price advantage.
We also dig into xAI’s safety card. Several jailbreak and harmful-compliance measurements improved, while general harmful-request compliance and the self-harm metric moved slightly in the wrong direction. Some biosecurity-related capability scores also declined intentionally, according to xAI, as part of its safety training approach.
And there are still major unanswered questions: no parameter count, no disclosed dense-vs-MoE architecture, no training compute or cost, no standalone long-context benchmark, and no general computer-use tool.
Grok 4.7 is clearly more than a rename. The question is whether its gains justify how much more inference it may need to produce them.
CHAPTERS
00:00 Grok 4.7’s Huge Benchmark Jump Has a Catch
01:12 What Actually Shipped, and Where You Can’t Use It Yet
02:42 What xAI Actually Changed Under the Hood
04:22 The Coding Benchmarks, One at a Time
06:58 Agents That Work for Hours, and the Push Into Professional Work
09:06 The Independent Reality Check
11:09 The Price Tag, and the Fine Print xAI Didn’t Put in the Headline
13:00 What Got Safer, and What Quietly Got Worse
14:30 What xAI Still Won’t Tell You
#grok47 #xai #grok #ai #artificialintelligence



