Imagine two programmers facing a damp patch on a wall. One paints over it and closes the ticket. The other follows the stain to the leaking pipe, repairs the pipe, then checks the next room. Claude Opus 5 is built to be the second programmer.
That launch evaluation is now called Terminal-Bench 3.0. Anthropic originally reported 43.3%; the current public board shows about 42.7% for Opus 5 and 34.6% for the cited GPT-5.6 Sol configuration. The snapshots differ slightly, but both support the same lesson: Opus is unusually good at unfamiliar terminal work.
One example explains the personality behind the number. Asked to reconstruct a machine part from a drawing it could not directly view, Opus 5 wrote its own vision pipeline, extracted geometry from pixels, and built the part in FreeCAD. On another task it found the root cause of a package-manager bug and fixed an edge case the community patch missed. These are vendor examples, not independent laws of nature, but they show what Anthropic tuned: keep investigating until the evidence agrees.
Why it moves above Fable 5
Fable still wins a few finish-line photographs. It scores 53.5% on FrontierCode against Opus 5’s 53.4%, and its maximum CursorBench 3.2 score is about half a point higher. But a ranking is a buying decision, not a museum of decimals. Opus 5 charges $5/$25 per million input/output tokens, exactly half Fable’s $10/$50, works without Fable’s general data-retention requirement, and has less restrictive classifiers. When performance is almost tied, price and deployability are part of performance.
The independent signal has strengthened. Artificial Analysis v4.1.1 gives Opus 5 roughly 63 and first place on its Intelligence Index, while its Coding Agent Index is about 78.0. The same testing keeps the warning label: max effort can generate a great many tokens.
Why it narrowly misses #1
GPT-6 Astra now leads DeepSWE v1.1 at 74.1%, just ahead of Opus 5 at 73.7% in the September tables, plus Terminal-Bench 4.0 at 57.9% and state-of-the-art computer use. Artificial Analysis re-scaled its coding-agent board at Astra’s launch (Fable 5.1 about 70, Astra 67), so the older Opus-era numbers are not comparable cells. Our #1 still goes to Astra on a Codex, efficiency, and execution-value tiebreak—not because it plainly outcodes Opus.
So route by job:
| Job | Best starting point |
|---|---|
| Vague bug, root-cause hunt, multi-file change, visual frontend verification | Opus 5 |
| Terminal-heavy work where speed, cost per finished task, and execution matter most | GPT-6 Astra |
| A peak CursorBench/FrontierCode task where budget is secondary | Fable 5 |
| Normal tickets with a firm cost ceiling | Use lower effort first, then escalate |
The practical upgrade path is unusually kind. Opus 5 keeps the 1M context window, 128k maximum output, vision, PDFs, prompt caching, batch processing, and Claude’s tools. Thinking is adaptive and on by default; effort runs from low through max. Existing Opus 4.8 prompts should mostly transfer, but developers must note that API web fetch and Priority Tier are unavailable at launch.
The bottom line is simple: Opus 5 replaces Opus 4.8 completely and makes Fable 5 harder to justify for routine production volume. Fable remains the expensive specialist with a few narrow leads. Opus 5 is the model you can give to more engineers, on more tasks, for longer—and still afford the verification loop that makes its judgment valuable.