Ranked #4 Everyday Ecosystem — The Leading AI Assistants
Anthropic

Claude Sonnet 5.5

Claude Sonnet 5.5 is the assistant for the ordinary working week: summarizing documents, building spreadsheets and slides, drafting reports. On independent office-work tests it lands within a couple of points of Opus 5.5, it's available on Claude's free plan, and its API tokens cost half as much. Opus still knows more facts and handles the truly messy problems better.

Updated September 29, 2026 Knowledge WorkFreemium1M Context
9.7out of 10
Official Website
Best for

Claude Sonnet 5.5 is the assistant for the ordinary working week: summarizing documents, building spreadsheets and slides, drafting reports. On independent office-work tests it lands within a couple of points of Opus 5.5, it's available on Claude's free plan, and its API tokens cost half as much. Opus still knows more facts and handles the truly messy problems better.

Why It Wins

Scores 56 on the independent Artificial Analysis Intelligence Index at max effort, two points behind Opus 5.5 (58). Artificial Analysis also measured 1844 Elo on GDPval-AA v2.1 (Opus 5.5: 1846) and 1811 on AA-Briefcase (Opus 5.5: 1822). It hallucinates less often than Opus 5.5 when it doesn't know an answer. API tokens cost $2/$10, and output is 30%+ faster than Sonnet 5.

Watch out

It knows fewer facts than Opus 5.5 (54% vs 66% accuracy on Artificial Analysis's knowledge test), so double-check niche details. At max effort it produces more text than any model Artificial Analysis has measured, which makes max expensive and slow. There is no image, audio, or video generation, and higher-risk security requests fall back to the older Sonnet 5.

01

What It Actually Is

Imagine an office with two senior people. One is the partner you call into the meeting when the deal is unusual and the stakes are high. The other is the colleague who gets through the week’s pile, forty emails, three spreadsheets, a slide deck for Thursday, quickly and neatly.

Most of us need the second person far more often than the first. Claude Sonnet 5.5 is that colleague.

Anthropic released it on September 28, 2026, a week after Claude Opus 5.5, as the second model in the Claude 5.5 family. It is meant as the everyday complement to Opus: in Anthropic’s words, Opus is built for complex work requiring careful judgment, while Sonnet 5.5 is strongest at well-scoped everyday tasks, fixing bugs, and creating polished documents, slides, and spreadsheets.

How close is it to Opus?

Closer than you might expect, at least on the independent tests.

Artificial Analysis runs two tests built around finished professional work rather than quiz questions. In GDPval-AA v2.1, models produce real deliverables, such as a financial model, a legal memo, or a briefing, and the results are compared head-to-head like chess games. Sonnet 5.5 scored 1844 Elo. Opus 5.5 scored 1846, a gap too small to matter. On the newer AA-Briefcase, which tests longer multi-step knowledge work, it scored 1811 against Opus’s 1822.

The broadest independent measure is the Artificial Analysis Intelligence Index, a composite of hard evaluations. At max effort, Sonnet 5.5 scored 56, second only to Opus 5.5 at 58.

That is a huge step for a Sonnet. Its predecessor, Sonnet 5, scored 1449 on GDPval-AA. In practical terms, office work that used to need the flagship can now go to the cheaper, faster model.

Where Opus still pulls ahead

There are two places where the gap is real.

Facts. Artificial Analysis’s AA-Omniscience test asks thousands of factual questions. Sonnet 5.5 answered 54% correctly; Opus 5.5 answered 66%. A smaller model simply holds less knowledge. There is a silver lining: when Sonnet 5.5 didn’t know an answer, it made one up 47% of the time, compared with 59% for Opus 5.5. It knows less, but it is more willing to say so. Still, for an obscure date, a name, or a statistic, ask it to search the web or cite a source.

Judgment. Anthropic is unusually direct about this. On several tests, Sonnet 5.5 at max effort performs comparably to Opus 5.5. But in Anthropic’s own testing and that of outside testers, Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment. If a question has no obvious frame, such as a strategy decision or a contract with unusual terms, that is still Opus territory.

What it’s good at

Anthropic highlights documents, summaries, spreadsheets, and slides, and says the model has a sharp eye for design: it can follow a slide template closely enough that little editing is needed. Its chart reading improved dramatically. On Anthropic’s Chartography test, which asks models to read values from charts without tools, it scored 61.6%, up from 15.6% for Sonnet 5.

It is also quick. Anthropic says output arrives more than 30% faster than Sonnet 5, and early testers noticed: Zendesk says support tickets were processed 20% faster, and Box measured 2.4 times the speed of the previous version.

It can operate a computer, too. On OSWorld 2.1, where a model clicks and types its way through real desktop apps, Anthropic reports 80.1% under partial-credit scoring, close to Opus 5.5’s 81.8%.

Leave the dial in the middle

Sonnet 5.5 has five effort settings: low, medium, high, xhigh, and max. The Claude apps use medium by default. Here is what Artificial Analysis measured at each:

Effort Intelligence Index Cost per task
Low 36 $0.41
Medium 41 $0.59
High 47 $1.08
Xhigh 52 $2.74
Max 56 $7.60

The last step is the expensive one. At max, Sonnet 5.5 wrote about 193,000 output tokens per task, the most Artificial Analysis has ever measured, and about 60% more than Opus 5.5 at max. For most people in the Claude app this simply means waiting longer. For developers paying per token, it means that a maxed-out Sonnet is no longer the cheap option. If a task needs that much thinking, it usually deserves Opus.

The honest catch

No pictures. Claude reads screenshots, charts, and scanned PDFs well, but it doesn’t generate images, audio, or video. If you want one app that also draws and talks, ChatGPT and Gemini are broader.

Price per token is not price per task. Tokens cost the same as Sonnet 5, $2 per million in and $10 per million out. Anthropic says typical tasks cost up to 30% less because the model finishes in fewer steps. That holds at normal settings; at max, as the table shows, the bill climbs fast.

Some requests are guarded. Sonnet 5.5 is the first Sonnet to ship with the stricter cybersecurity safeguards Anthropic uses on its top models. Higher-risk security requests visibly fall back to Sonnet 5. For almost everyone, this never comes up.

Who should use it

If your day looks like… Reach for Why
Emails, summaries, spreadsheets, slide decks Claude Sonnet 5.5 Near-Opus office work, fast, on the free plan
High-stakes judgment, obscure facts, messy research Claude Opus 5.5 More knowledge and clearly stronger judgment
“Operate this software on my computer and finish the job” GPT-6 Astra Stronger, more economical computer-use agent
One app for chat, images, and voice ChatGPT or Gemini Broader creative toolbox

The everyday verdict

Claude Sonnet 5.5 won’t replace Opus for the hardest questions, and Anthropic doesn’t pretend it will. But the hardest questions are a small part of most weeks. For everything else, it is fast, capable, available on Claude’s free plan, and priced so you don’t have to think about it. Leave it on the default setting, check the obscure facts, and call in Opus when the stakes go up.

02

Strengths and honest limitations

Key Strengths

  • Close to the flagship on real work: Artificial Analysis ran two tests of finished professional deliverables. Sonnet 5.5 scored 1844 Elo on GDPval-AA v2.1 against 1846 for Opus 5.5, and 1811 on the long-horizon AA-Briefcase against 1822. For everyday office work, the two are nearly level.
  • Second on the broadest independent index: At max effort it scores 56 on the Artificial Analysis Intelligence Index, behind only Opus 5.5 at 58. Even at the everyday High setting it reaches 47 for about a dollar per task.
  • Says ‘I don’t know’ more often: On the AA-Omniscience knowledge test, Sonnet 5.5 hallucinated on 47% of the questions it couldn’t answer correctly, against 59% for Opus 5.5 — it is more willing to admit uncertainty.
  • Built for documents, slides, and spreadsheets: Anthropic says it is strongest at well-scoped everyday tasks and polished documents, with a sharp eye for design. It also reads charts far better than Sonnet 5: 61.6% on Anthropic’s Chartography test versus 15.6%.
  • Fast, widely available, and cheaper to build on: Output is 30%+ faster than Sonnet 5. Anyone can use it on Claude’s web and mobile apps, including the free plan, and developers pay $2/$10 per million tokens — half of Opus 5.5 — on Anthropic’s API, AWS, Google Cloud, and Microsoft Azure.

Honest Limitations

  • Fewer facts in its head: On AA-Omniscience it answered 54% of knowledge questions correctly, against 66% for Opus 5.5. For obscure names, dates, or figures, ask it to search or cite a source.
  • Max effort is not worth it: Artificial Analysis measured about 193,000 output tokens per task at max, the most it has recorded, and the cost per task nearly triples from xhigh to max. Everyday work belongs on the default settings.
  • Opus still wins the hard cases: Anthropic says Opus 5.5 remains clearly stronger at complex, open-ended work that needs sustained judgment — the ambiguous contract, the strategy question with no clear frame.
  • No pictures or voice studio: Claude reads images and charts but does not create images, audio, or video. ChatGPT and Gemini remain broader all-in-one toolboxes.
  • Some security requests are rerouted: Sonnet 5.5 ships with Anthropic’s stricter cyber safeguards, and higher-risk security tasks visibly fall back to Sonnet 5. Most people will never notice.
03

Benchmark Snapshot

Artificial Analysis Intelligence Index — 56 (max effort)

Independent composite of hard evaluations, measured September 2026. Second only to Opus 5.5 (58) at launch. By effort: low 36, medium 41, high 47, xhigh 52, max 56.

GDPval-AA v2.1 — 1844 Elo

Real professional deliverables across many occupations, run by Artificial Analysis. Opus 5.5 scores 1846 and Sonnet 5 1449.

AA-Briefcase v1.1 — 1811 Elo

Artificial Analysis's newer test of long, multi-step knowledge work. Opus 5.5 scores 1822 and Sonnet 5 1359.

AA-Omniscience — 54% accuracy, 47% hallucination rate

Independent knowledge test. Opus 5.5 answers more questions correctly (66%) but hallucinates more often when wrong (59%). Lower hallucination rate is better.

OSWorld 2.1 — 80.1% (partial reward)

Anthropic's figure for operating a real desktop. Opus 5.5 scores 81.8% and Sonnet 5 57.0% under the same partial-credit scoring.

Price — $2 / $10 per 1M tokens

Unchanged from Sonnet 5 and half of Opus 5.5. Artificial Analysis measured cost per task from $0.41 at low effort to $7.60 at max.

04

The Verdict

Claude Sonnet 5.5 is the assistant most people can use every day without thinking about cost: close to Opus 5.5 on independent office-work tests, on Claude’s free plan, and at half the API price. It ranks below Opus 5.5 because it knows noticeably fewer facts and, by Anthropic’s own account, handles open-ended judgment less well; it sits below GPT-6 Astra and Fable 5.1 because they lead on computer use and long research. For summaries, spreadsheets, slides, and reports, it is the sensible default.

05

Frequently Asked Questions