There’s a difference between someone who writes code and someone who understands a system.
The first person sees an error that says undefined is not a function, wraps that line in an if check, and moves on. The crash goes away. Three days later something else breaks, somewhere nobody expected.
The second person asks why the value was undefined. They follow it back through the services, find that a database write finished before a message was confirmed, and fix the order so the bad state can’t happen at all.
Claude Opus 5.5 is very much the second kind. Released on September 22, 2026, it sits at #2 on our Coding leaderboard, closer to #1 than any model before it.
What it’s good at: big, tangled jobs
The stories from Anthropic’s early testers all have the same shape: large jobs that used to take a team weeks.
- One tester finished a 680,000-line code migration in less than a day.
- Another audited and fixed a 200,000-line codebase in under three hours. Opus 5 needed more than 20 hours and 2.5 times as many tokens for the same job.
- In an internal test, Opus 5.5 and Fable 5.1 both rewrote HAProxy, the widely used web traffic balancer, from C into Rust, and both passed nearly all of HAProxy’s own tests. Opus 5.5 finished in 9.5 hours instead of 12, at 51% lower cost.
These are vendor anecdotes, not independent measurements, so treat them as signs of direction rather than guarantees. But the pattern is consistent. The larger and messier the job, the more Opus 5.5 pulls ahead.
It also explains itself clearly. Anthropic’s launch post shows it diagnosing a billing bug. Instead of a wall of technical detail, it opens with the money: $1.50 of a customer’s drop came from a pricing change, and the other $9.92 came from a bug in a specific commit that stopped counting usage on the last day of each month. When an agent is changing your code, reports like that are how you keep it honest.
The benchmarks: a tie, depending on who’s counting
This is where it is worth slowing down, because vendor tables and independent tests tell slightly different stories.
Anthropic’s numbers (September 2026):
- Terminal-Bench 4.0, which tests multi-step work in a command-line terminal: Opus 5.5 scored 66.4% (±2.6 points). GPT-6 Astra scored 57.9%, a figure OpenAI reported.
- FrontierCode v1.1, which checks whether code changes are good enough to merge: Opus 5.5 54.4%, Astra 53.3%.
- CursorBench 4.0, which covers longer coding tasks inside the Cursor editor: Opus 5.5 57.8%, Fable 5.1 51.8%.
Independent numbers: Artificial Analysis ran Terminal-Bench 4.0 itself and got 59.6% for both Opus 5.5 and Astra. That’s a dead heat.
Our rule is to trust independent measurements over vendor tables. That gives us a tie on raw ability. So the ranking comes down to something else.
Why GPT-6 Astra keeps #1
When two models are equally good, we rank them by what a finished task costs you.
Here Astra has a clear edge. At max effort, Artificial Analysis measured Opus 5.5 writing roughly 119,000 output tokens per task, against about 27,000 for Astra. Opus’s tokens are cheaper ($20 per million output against Astra’s $50), but it uses about four times as many. Do the arithmetic and the average Astra task at peak performance costs roughly half as much.
Anthropic makes a fair counterpoint. At default effort, it says Opus 5.5 matches Astra on Terminal-Bench for about 40% of the cost, and beats it on FrontierCode at about 20% of the cost. If independent testing confirms that, the ranking will change. For now it’s a vendor claim, and we wait for independent data before promoting a model.
The honest limitations
- Keep max effort for hard problems. Left on max for routine fixes, Opus 5.5 will write long reasoning traces you pay for.
- Security work is rerouted. Opus 5.5 is strong enough at cybersecurity that Anthropic sends most security tasks to Opus 4.8 unless you’re in its Cyber Verification Program. Normal bug fixing isn’t affected.
- Long sessions still hit limits. Anthropic raised the five-hour limits on paid plans, but a multi-hour agent run over a large repository can still bump into them.
Which one to hire
Use GPT-6 Astra when you have a queue of well-defined tickets and want them closed as cheaply as possible.
Use Claude Opus 5.5 when the work is big, ambiguous, or risky: a framework migration, a deep audit, or a bug that crosses service boundaries. On those jobs, careful reasoning is worth more than terse output. It is the engineer we would trust with the system nobody else wants to touch.