July 25 update: Opus 5 moves Fable 5 to #3 in our coding ranking. Fable still owns the published SWE-Bench Pro result and tiny CursorBench/FrontierCode edges, but Opus 5 wins Frontier-Bench and delivers near-parity at half the token price. The June launch review below remains useful evidence, not the current overall order.
There’s a number that makes this review easy to write: 80.3%. That’s Claude Fable 5 on SWE-Bench Pro — the benchmark that doesn’t care about toy problems, only whether AI can fix actual bugs in actual production codebases. GPT-5.5 scores 58.6%. The previous king, Opus 4.8, scored 69.2%. Fable 5 doesn’t just win — it wins by a margin that makes you double-check the numbers.
But SWE-Bench Pro is only half the story. FrontierCode Diamond — Cognition’s benchmark for whether models can write token-efficient, production-quality code — tells the other half. Fable 5: 29.3%. Opus 4.8: 13.4%. GPT-5.5: 5.7%. That’s not a lead; that’s a different sport. And the model achieves these scores at medium reasoning effort, meaning it burns fewer tokens to produce better code. The expensive model that’s actually cheaper per real-world task.
The Stripe case study isn’t a press release fantasy. A 50-million-line Ruby codebase — the kind of monolith that makes engineers sweat — got migrated in a single day. Work that would have taken a full team two months. The model planned, executed, self-verified, and delivered. On CursorBench, Cursor’s CEO said it “opened up a class of long-horizon problems that were out of reach for earlier models.” On the Senior Engineer Benchmark, it scored 91/100 — while GPT-5.5 and Opus 4.8 both landed in the low 60s.
This is what Mythos-class architecture looks like when you wrap it in safety guardrails and hand it to developers. Following temporary export control reviews in mid-June, Fable 5 was fully restored globally on June 30, 2026. Shortly after, Amazon security engineers credited Fable 5 with autonomously identifying and patching a zero-day authentication vulnerability across 12,000 internal services in hours. The guardrails are real — queries on cybersecurity exploit generation, biology, and chemistry get routed to Opus 4.8 via a specialized classifier. But for the 95%+ of software engineering work that doesn’t trigger safety classifiers, you’re working with the most capable model ever released to the public. The agentic coding era just got its clearest champion.