Imagine two programmers facing a damp patch on a wall. One paints over it and closes the ticket. The other follows the stain to the leaking pipe, repairs the pipe, then checks the next room. Claude Opus 5 is built to be the second programmer.
That is why the most important launch number is 43.3% on Frontier-Bench v0.1. The evaluation asks an agent to solve unfamiliar engineering jobs through a terminal and tools. In Anthropic’s comparison, GPT-5.6 Sol scored 34.4%, Fable 5 scored 33.7%, and Opus 4.8 scored 21.1%. Opus 5 did not merely replace 4.8; it roughly doubled it.
One example explains the personality behind the number. Asked to reconstruct a machine part from a drawing it could not directly view, Opus 5 wrote its own vision pipeline, extracted geometry from pixels, and built the part in FreeCAD. On another task it found the root cause of a package-manager bug and fixed an edge case the community patch missed. These are vendor examples, not independent laws of nature, but they show what Anthropic tuned: keep investigating until the evidence agrees.
Why it moves above Fable 5
Fable still wins a few finish-line photographs. It scores 53.5% on FrontierCode against Opus 5’s 53.4%, and its maximum CursorBench 3.2 score is about half a point higher. But a ranking is a buying decision, not a museum of decimals. Opus 5 charges $5/$25 per million input/output tokens, exactly half Fable’s $10/$50, works without Fable’s general data-retention requirement, and has less restrictive classifiers. When performance is almost tied, price and deployability are part of performance.
The independent signal arrived quickly. Artificial Analysis gives Opus 5 max effort 61 and first place on its Intelligence Index. High and xhigh score 59 and 60, which suggests the model scales sensibly with extra thought. The same testing also shows the warning label: max effort generated a great many tokens. A half-price model can still produce a full-price invoice if allowed to write a novel while fixing a button.
Why it narrowly misses #1
GPT-5.6 Sol leads DeepSWE v1.1 at 72.7%, ahead of Fable at 69.7% and Opus 5 at 68.8%. In Artificial Analysis’ current direct comparison, however, Claude Code with Opus 5 and Codex with GPT-5.6 Sol tie at 67 on the overall Coding Agent Index. The split is more useful than the tie: GPT wins DeepSWE and Terminal-Bench and finishes tasks faster and more cheaply, while Opus wins the repository-understanding test. Our #1 slot therefore goes to GPT-5.6 on an efficiency-and-terminal tiebreak, not because it plainly outcodes Opus 5 everywhere.
So route by job:
| Job | Best starting point |
|---|---|
| Vague bug, root-cause hunt, multi-file change, visual frontend verification | Opus 5 |
| Terminal-heavy work where speed, cost per finished task, and execution matter most | GPT-5.6 Sol |
| A peak CursorBench/FrontierCode task where budget is secondary | Fable 5 |
| Normal tickets with a firm cost ceiling | Use lower effort first, then escalate |
The practical upgrade path is unusually kind. Opus 5 keeps the 1M context window, 128k maximum output, vision, PDFs, prompt caching, batch processing, and Claude’s tools. Thinking is adaptive and on by default; effort runs from low through max. Existing Opus 4.8 prompts should mostly transfer, but developers must note that API web fetch and Priority Tier are unavailable at launch.
The bottom line is simple: Opus 5 replaces Opus 4.8 completely and makes Fable 5 harder to justify for routine production volume. Fable remains the expensive specialist with a few narrow leads. Opus 5 is the model you can give to more engineers, on more tasks, for longer—and still afford the verification loop that makes its judgment valuable.