Ranked #2 Coding — AI That Writes Production Code
Anthropic

Claude Opus 5.5

Claude Opus 5.5 is the coding model to call when the job is big and tangled: a migration across hundreds of thousands of lines, a bug that crosses five services, a legacy system nobody fully understands anymore. On Anthropic's own benchmarks it leads GPT-6 Astra; on independent testing the two are tied. It holds our #2 spot only because Astra finishes the same terminal work in far fewer tokens.

Updated September 27, 2026 Agentic CodingTerminal-Bench 4.0FrontierCode 54.4%
9.9out of 10
Official Website
Best for

Claude Opus 5.5 is the coding model to call when the job is big and tangled: a migration across hundreds of thousands of lines, a bug that crosses five services, a legacy system nobody fully understands anymore. On Anthropic's own benchmarks it leads GPT-6 Astra; on independent testing the two are tied. It holds our #2 spot only because Astra finishes the same terminal work in far fewer tokens.

Why It Wins

Anthropic reports 66.4% on Terminal-Bench 4.0, 54.4% on FrontierCode v1.1, and 57.8% on CursorBench 4.0 — ahead of GPT-6 Astra on the first two. Independent Artificial Analysis testing puts Opus 5.5 and Astra in a dead heat on Terminal-Bench 4.0 at 59.6%. Early testers report a 680,000-line migration in under a day and a 200,000-line audit in under three hours. Tokens cost $4/$20 and cache reads $0.20.

Watch out

At max effort Opus 5.5 is wordy — Artificial Analysis measured about four times the output tokens of GPT-6 Astra per task — which is why Astra keeps our #1 on cost per finished job. Most cybersecurity work is rerouted to Opus 4.8 by Anthropic's safeguards, and heavy sessions can still hit plan usage limits.

01

What It Actually Is

There’s a difference between someone who writes code and someone who understands a system.

The first person sees an error that says undefined is not a function, wraps that line in an if check, and moves on. The crash goes away. Three days later something else breaks, somewhere nobody expected.

The second person asks why the value was undefined. They follow it back through the services, find that a database write finished before a message was confirmed, and fix the order so the bad state can’t happen at all.

Claude Opus 5.5 is very much the second kind. Released on September 22, 2026, it sits at #2 on our Coding leaderboard, closer to #1 than any model before it.

What it’s good at: big, tangled jobs

The stories from Anthropic’s early testers all have the same shape: large jobs that used to take a team weeks.

  • One tester finished a 680,000-line code migration in less than a day.
  • Another audited and fixed a 200,000-line codebase in under three hours. Opus 5 needed more than 20 hours and 2.5 times as many tokens for the same job.
  • In an internal test, Opus 5.5 and Fable 5.1 both rewrote HAProxy, the widely used web traffic balancer, from C into Rust, and both passed nearly all of HAProxy’s own tests. Opus 5.5 finished in 9.5 hours instead of 12, at 51% lower cost.

These are vendor anecdotes, not independent measurements, so treat them as signs of direction rather than guarantees. But the pattern is consistent. The larger and messier the job, the more Opus 5.5 pulls ahead.

It also explains itself clearly. Anthropic’s launch post shows it diagnosing a billing bug. Instead of a wall of technical detail, it opens with the money: $1.50 of a customer’s drop came from a pricing change, and the other $9.92 came from a bug in a specific commit that stopped counting usage on the last day of each month. When an agent is changing your code, reports like that are how you keep it honest.

The benchmarks: a tie, depending on who’s counting

This is where it is worth slowing down, because vendor tables and independent tests tell slightly different stories.

Anthropic’s numbers (September 2026):

  • Terminal-Bench 4.0, which tests multi-step work in a command-line terminal: Opus 5.5 scored 66.4% (±2.6 points). GPT-6 Astra scored 57.9%, a figure OpenAI reported.
  • FrontierCode v1.1, which checks whether code changes are good enough to merge: Opus 5.5 54.4%, Astra 53.3%.
  • CursorBench 4.0, which covers longer coding tasks inside the Cursor editor: Opus 5.5 57.8%, Fable 5.1 51.8%.

Independent numbers: Artificial Analysis ran Terminal-Bench 4.0 itself and got 59.6% for both Opus 5.5 and Astra. That’s a dead heat.

Our rule is to trust independent measurements over vendor tables. That gives us a tie on raw ability. So the ranking comes down to something else.

Why GPT-6 Astra keeps #1

When two models are equally good, we rank them by what a finished task costs you.

Here Astra has a clear edge. At max effort, Artificial Analysis measured Opus 5.5 writing roughly 119,000 output tokens per task, against about 27,000 for Astra. Opus’s tokens are cheaper ($20 per million output against Astra’s $50), but it uses about four times as many. Do the arithmetic and the average Astra task at peak performance costs roughly half as much.

Anthropic makes a fair counterpoint. At default effort, it says Opus 5.5 matches Astra on Terminal-Bench for about 40% of the cost, and beats it on FrontierCode at about 20% of the cost. If independent testing confirms that, the ranking will change. For now it’s a vendor claim, and we wait for independent data before promoting a model.

The honest limitations

  • Keep max effort for hard problems. Left on max for routine fixes, Opus 5.5 will write long reasoning traces you pay for.
  • Security work is rerouted. Opus 5.5 is strong enough at cybersecurity that Anthropic sends most security tasks to Opus 4.8 unless you’re in its Cyber Verification Program. Normal bug fixing isn’t affected.
  • Long sessions still hit limits. Anthropic raised the five-hour limits on paid plans, but a multi-hour agent run over a large repository can still bump into them.

Which one to hire

Use GPT-6 Astra when you have a queue of well-defined tickets and want them closed as cheaply as possible.

Use Claude Opus 5.5 when the work is big, ambiguous, or risky: a framework migration, a deep audit, or a bug that crosses service boundaries. On those jobs, careful reasoning is worth more than terse output. It is the engineer we would trust with the system nobody else wants to touch.

02

Strengths and honest limitations

Key Strengths

  • Built for the big jobs: Anthropic’s testers used it for a 680,000-line code migration finished in under a day, and a 200,000-line audit in under three hours where Opus 5 needed over 20 hours and 2.5 times the tokens. Large, sprawling work is where it pulls ahead.
  • Finds the real bug: It traces a failure back to where it started instead of wrapping the symptom in a try/catch. In Anthropic’s example, it explained a billing bug in plain language: what broke, which commit broke it, and how much money was lost.
  • Merge-ready changes: FrontierCode v1.1 checks whether a model’s code changes are good enough to merge into real projects. Opus 5.5 scores 54.4% in Anthropic’s table, narrowly ahead of GPT-6 Astra at 53.3%.
  • Reports you can actually read: Anthropic retrained how the model writes. Explanations lead with the conclusion and skip the jargon, which makes reviewing an agent’s work much faster.
  • Harder to trick: Anthropic says Opus 5.5 ties Fable 5.1 for the lowest prompt-injection success rate in Gray Swan’s testing, and it is much less likely than earlier models to take hard-to-reverse actions. That matters when an agent runs unattended in your repository.

Honest Limitations

  • Wordy at max effort: Artificial Analysis measured roughly 119,000 output tokens per task at max effort, against about 27,000 for GPT-6 Astra. At the same peak score, that makes Astra the cheaper model per finished task despite its higher sticker price.
  • Vendor numbers still dominate: The headline 66.4% Terminal-Bench score is Anthropic’s own run (standard error ±2.6 points), and Astra’s 57.9% comparison figure comes from OpenAI. The independent Artificial Analysis run shows a tie, not a lead.
  • Security work is rerouted: Because Opus 5.5 is very capable at cybersecurity, most security tasks are passed to Opus 4.8 unless you join Anthropic’s Cyber Verification Program. Ordinary bug fixing is not affected.
  • Usage limits on long sessions: Anthropic raised five-hour limits on paid plans, but multi-hour agent runs over large codebases can still hit them.
03

Benchmark Snapshot

Terminal-Bench 4.0 — 59.6% independent tie with Astra (66.4% vendor)

Complex multi-step tasks in a command-line terminal. Artificial Analysis measured Opus 5.5 and GPT-6 Astra (xhigh) at 59.6% each. Anthropic's own run gives Opus 5.5 66.4% (±2.6) versus 57.9% for Astra as reported by OpenAI.

FrontierCode v1.1 — 54.4%

Whether code changes are ready to merge into real codebases. Anthropic's launch table: Opus 5.5 54.4%, GPT-6 Astra 53.3%, Fable 5.1 50.3%, Opus 5 48.0%.

CursorBench 4.0 — 57.8%

Longer-running coding tasks inside the Cursor editor. Anthropic reports 57.8% against Fable 5.1's 51.8% and GPT-5.6 Sol's 41.7%.

Output tokens per task — ~119k at max effort

Artificial Analysis measurement at max effort, versus ~27k for GPT-6 Astra. This is the reason Astra keeps our #1: the same peak score for roughly half the cost per task.

Price — $4 / $20 per 1M tokens, $0.20 cache read

Cache reads, which dominate coding-agent bills, dropped 60% from Opus 5. Fast mode (up to 2.5x speed) costs $8/$40.

04

The Verdict

Claude Opus 5.5 is our #2 coding model, and the gap to #1 is the narrowest it has ever been. On independent Terminal-Bench testing it ties GPT-6 Astra; on Anthropic’s own FrontierCode and Terminal-Bench runs it leads. We keep Astra on top for one declared reason: at peak performance, Astra finishes the same work in about a quarter of the output tokens, so each completed task costs less. If your work is a sprawling migration, a deep audit, or a bug nobody else can find, Opus 5.5 is the model we would trust with it — and at default effort, it is cheaper than ever.

05

Frequently Asked Questions