Ranked #2 Coding — AI That Writes Production Code
Anthropic

Claude Opus 5

The evidence-leading frontier coder with unusually patient verification. Opus 5 remains our #2 only because this guide gives GPT-6 Astra an explicit tiebreak for Codex integration, terminal and computer-use efficiency, and cost per finished task.

Updated August 29, 2026 Agentic CodingFrontier-Bench SOTARoot-Cause Debugging
9.9out of 10
Official Website
Best for

The evidence-leading frontier coder with unusually patient verification. Opus 5 remains our #2 only because this guide gives GPT-6 Astra an explicit tiebreak for Codex integration, terminal and computer-use efficiency, and cost per finished task.

Why It Wins

Launch-period Artificial Analysis coding-agent evidence put Opus around 78.0, narrowly above GPT-5.6 Sol at 77.4; the board was re-scaled at GPT-6 Astra's September launch (Fable 5.1 about 70, Astra 67), so those generations of numbers are not comparable. The public Terminal-Bench 3.0 board records about 42.7%, while Intelligence Index v4.1.1 is about 63. It retains 1M context, 128K output, and $5/$25 pricing.

Watch out

Its #2 rank is an editorial product judgment, not the live benchmark order. GPT-6 Astra still leads Terminal-Bench, DeepSWE, and the science and SRE lanes, while Opus max effort can be verbose and slower. Anthropic's launch scores and the current public boards are related snapshots, not interchangeable numbers.

01

What It Actually Is

Imagine two programmers facing a damp patch on a wall. One paints over it and closes the ticket. The other follows the stain to the leaking pipe, repairs the pipe, then checks the next room. Claude Opus 5 is built to be the second programmer.

That launch evaluation is now called Terminal-Bench 3.0. Anthropic originally reported 43.3%; the current public board shows about 42.7% for Opus 5 and 34.6% for the cited GPT-5.6 Sol configuration. The snapshots differ slightly, but both support the same lesson: Opus is unusually good at unfamiliar terminal work.

One example explains the personality behind the number. Asked to reconstruct a machine part from a drawing it could not directly view, Opus 5 wrote its own vision pipeline, extracted geometry from pixels, and built the part in FreeCAD. On another task it found the root cause of a package-manager bug and fixed an edge case the community patch missed. These are vendor examples, not independent laws of nature, but they show what Anthropic tuned: keep investigating until the evidence agrees.

Why it moves above Fable 5

Fable still wins a few finish-line photographs. It scores 53.5% on FrontierCode against Opus 5’s 53.4%, and its maximum CursorBench 3.2 score is about half a point higher. But a ranking is a buying decision, not a museum of decimals. Opus 5 charges $5/$25 per million input/output tokens, exactly half Fable’s $10/$50, works without Fable’s general data-retention requirement, and has less restrictive classifiers. When performance is almost tied, price and deployability are part of performance.

The independent signal has strengthened. Artificial Analysis v4.1.1 gives Opus 5 roughly 63 and first place on its Intelligence Index, while its Coding Agent Index is about 78.0. The same testing keeps the warning label: max effort can generate a great many tokens.

Why it narrowly misses #1

GPT-6 Astra now leads DeepSWE v1.1 at 74.1%, just ahead of Opus 5 at 73.7% in the September tables, plus Terminal-Bench 4.0 at 57.9% and state-of-the-art computer use. Artificial Analysis re-scaled its coding-agent board at Astra’s launch (Fable 5.1 about 70, Astra 67), so the older Opus-era numbers are not comparable cells. Our #1 still goes to Astra on a Codex, efficiency, and execution-value tiebreak—not because it plainly outcodes Opus.

So route by job:

Job Best starting point
Vague bug, root-cause hunt, multi-file change, visual frontend verification Opus 5
Terminal-heavy work where speed, cost per finished task, and execution matter most GPT-6 Astra
A peak CursorBench/FrontierCode task where budget is secondary Fable 5
Normal tickets with a firm cost ceiling Use lower effort first, then escalate

The practical upgrade path is unusually kind. Opus 5 keeps the 1M context window, 128k maximum output, vision, PDFs, prompt caching, batch processing, and Claude’s tools. Thinking is adaptive and on by default; effort runs from low through max. Existing Opus 4.8 prompts should mostly transfer, but developers must note that API web fetch and Priority Tier are unavailable at launch.

The bottom line is simple: Opus 5 replaces Opus 4.8 completely and makes Fable 5 harder to justify for routine production volume. Fable remains the expensive specialist with a few narrow leads. Opus 5 is the model you can give to more engineers, on more tasks, for longer—and still afford the verification loop that makes its judgment valuable.

02

Strengths and honest limitations

Key Strengths

  • Terminal-Bench 3.0 leadership: The benchmark formerly introduced as Frontier-Bench now lists Opus 5 around 42.7% on its public board, ahead of GPT-5.6 Sol’s cited 34.6% configuration. It rewards an agent that can live in a terminal and finish unfamiliar engineering work.
  • Verification before celebration: In Anthropic’s examples, Opus 5 found root causes that another model patched only at the surface, built its own test harness when no live data feed existed, and checked web interfaces at desktop and phone widths before declaring them finished.
  • Near-Fable coding at half the token price: At max effort it comes within 0.5 percentage points of Fable 5’s peak CursorBench 3.2 result, according to Anthropic, while input and output tokens cost $5/$25 instead of $10/$50.
  • A full repository can fit on the desk: The 1M-token context window is the default, max output is 128k, and Anthropic says instruction following, tool use, and reasoning remain consistent across long context. That matters for migrations and multi-file investigations.
  • Available where teams already build: The model ID is claude-opus-5; it ships in Claude Code and the Claude API, with listings for Amazon Bedrock, Google Cloud Vertex AI, and Microsoft Foundry. Fast mode trades twice the base price for roughly 2.5× speed.

Honest Limitations

  • GPT-6 Astra still owns important lanes: Opus 5 posts 73.7% on DeepSWE v1.1, just behind Astra’s 74.1% in the September tables, and Astra leads Terminal-Bench 4.0 at 57.9% while using roughly a third of Sol’s tokens. Artificial Analysis re-scaled its coding board at Astra’s launch (Fable 5.1 about 70, Astra 67), so cross-time index comparisons are not like-for-like — but Astra’s Codex integration and token efficiency remain our #1 tiebreak.
  • Fable still has narrow peak wins: Fable 5 scores 53.5% to Opus 5’s 53.4% on FrontierCode 1.1 Main, and Anthropic describes Opus 5 as just under Fable’s maximum CursorBench score. Small gaps are not universal truths, but neither should they be erased.
  • Max effort can eat the savings: Independent Artificial Analysis testing found Opus 5 max highly capable but verbose. A cheaper token does not help if an agent spends twice as many of them; sweep effort levels on your own tasks.
  • Migration has two platform catches: Web fetch and Priority Tier are not supported at launch. Thinking is on by default, non-default sampling parameters are rejected, and teams coming from older Claude integrations should retest token budgets and prompts.
  • Security boundaries remain visible: Opus 5 can find source-code vulnerabilities, but classifiers block penetration testing, exploit generation, and binary-based vulnerability scanning outside approved programs. Defensive specialists may still encounter false positives.
03

Benchmark Snapshot

Terminal-Bench 3.0 — about 42.7%

The current public result follows the earlier Frontier-Bench launch run of 43.3%. Snapshot and harness details differ, so use the current name and date the score.

CursorBench 3.2 — within 0.5 points of Fable 5

At max effort, Anthropic reports near-parity with Fable's peak at half the cost per task; Cursor warns that tiny score differences may not be statistically meaningful.

DeepSWE v1.1 — 68.8%

Strong long-horizon repository work, but behind Fable 5 at 69.7% and GPT-5.6 Sol at 72.7% in Anthropic's comparison.

FrontierCode 1.1 Main — 53.4%

Effectively tied with Fable 5 at 53.5% and ahead of GPT-5.6 Sol at 47.5%; the tenth of a point is not a practical moat.

Artificial Analysis Intelligence Index — about 63 (#1)

Current v4.1.1 independent evidence places Opus 5 at the broad intelligence frontier. The earlier 61 launch-week score is now stale.

04

The Verdict

Claude Opus 5 is our #2 coding model even though current broad independent evidence places it narrowly first. That is deliberate: this guide gives GPT-6 Astra a product tiebreak for Codex integration, terminal and computer-use efficiency, routing, and cost per finished task. Start Opus on vague, multi-file work where architecture and verification matter; start Astra on terminal-heavy throughput and desktop-driven work. The ranking exposes the weighting instead of rewriting the live leaderboard.

05

Frequently Asked Questions