Ranked #1 Everyday Ecosystem — The Leading AI Assistants
Anthropic

Claude — Opus 5

The new everyday intelligence leader: Opus 5 brings Fable-like judgment to a model people and companies can actually run at scale. It leads early independent intelligence testing, tops Anthropic's knowledge-work and computer-use comparisons, costs half as much as Fable 5, and is broadly available across Claude, cloud platforms, and the API.

Updated July 25, 2026 Knowledge Work SOTA1M ContextReasoning
9.9out of 10
Official Website
Best for

The new everyday intelligence leader: Opus 5 brings Fable-like judgment to a model people and companies can actually run at scale. It leads early independent intelligence testing, tops Anthropic's knowledge-work and computer-use comparisons, costs half as much as Fable 5, and is broadly available across Claude, cloud platforms, and the API.

Why It Wins

GDPval-AA v2 Elo 1861, ARC-AGI-3 30.2%, BrowseComp 90.8%, OSWorld 2.0 70.6%, AutomationBench 26.0%, and 64.7% on tool-assisted Humanity's Last Exam. Artificial Analysis independently scores max effort at 61 and #1 overall. 1M context, 128k output, $5/$25 pricing, and no general-access data-retention requirement.

Watch out

The crown is provisional because independent testing is less than 48 hours old. Fable 5 still wins some specialist and no-tools evaluations, GPT-5.6 remains the broader consumer ecosystem, and Opus 5 has no native image generation. Max effort is slow and verbose, while Claude's free and paid usage limits still matter.

01

What It Actually Is

Most chatbots can feel like a brilliant intern on a frantic Monday: quick, eager, and perfectly capable of presenting a polished spreadsheet whose grand total is quietly wrong. Claude Opus 5 feels more like the colleague who checks the formulas before pressing Send.

That judgment is the reason it moves straight to #1 in our everyday category. Anthropic’s launch table gives it 1861 Elo on GDPval-AA v2, a knowledge-work evaluation built around useful professional tasks. Fable 5 scores 1747; GPT-5.6 Sol, 1736; Opus 4.8, 1593. Box reports an 8% overall improvement over 4.8, with larger gains in data analysis and due diligence. A financial-evaluation customer saw nine points more accuracy with fewer turns and 60% less time. These are early customer reports, but they describe the same shape as the benchmarks: careful work with fewer trips around the block.

More than a well-read typewriter

Opus 5 scores 70.6% on OSWorld 2.0, which asks a model to use a computer, not merely talk about one. It reaches 26.0% on AutomationBench, roughly one and a half times the next result in Anthropic’s comparison. Zapier reports that Opus 5 completed an entire customer-retention workflow: it read account-health data, identified customers at risk of leaving, alerted the people responsible for those accounts, and prepared a summary. Earlier models failed to finish that chain.

The strangest number is 30.2% on ARC-AGI-3. This is an interactive test of solving unfamiliar problems, closer to learning the rules of a new board game than recalling a fact. GPT-5.6 Sol scored 7.8% in the same table. A benchmark is not a soul, but this one suggests Opus 5 is unusually good when the instructions stop holding its hand.

Independent testing supports the direction. Artificial Analysis gives Opus 5 max effort 61, currently first on its Intelligence Index. High scores 59 and xhigh 60. The ladder matters: you can buy more thought when a contract or financial model deserves it, then turn the dial down for ordinary drafting.

Why it ranks above Fable 5

Fable is still slightly better in a few narrow places: Humanity’s Last Exam without tools, FrontierCode, the highest CursorBench setting, and a held-out legal-agent test. If your private evaluation matches one of those lanes, use it. But Fable costs $10/$50 per million input/output tokens. Opus 5 costs $5/$25, has no general-access data-retention requirement, and its safety classifiers are expected to intervene much less often.

Think of Fable as a racing prototype and Opus 5 as the road car that is almost as fast, costs half as much to run, and can legally travel far beyond the track. For daily work, the road car is the more useful machine.

Why ChatGPT still belongs in the conversation

GPT-5.6 remains a superb all-round system. ChatGPT combines Work, Codex, image generation, Sites, connected apps, a browser, and desktop computer use in one enormous ecosystem. Claude cannot generate images natively, and its plan limits can be tighter. Our ranking says Opus 5 offers the best current blend of reasoning quality, professional agency, price, and deployability; it does not say every person should cancel ChatGPT.

The hardware behind that workbench is substantial: 1M tokens of context by default, 128k maximum output, vision for charts and documents, PDFs, files, memory, web search, tools, and computer use. The API model is claude-opus-5, with support across Anthropic’s platform and major clouds. Fast mode runs around 2.5 times faster at twice the base token price.

There are caveats. Independent data is still young. Max effort can be verbose, and xhigh has noticeable first-token latency. API web fetch and Priority Tier are absent at launch. The sensible habit is to start lower, verify the result, and raise effort only when the cost of being wrong is larger than the cost of thinking.

Opus 5 is not merely a cheaper Fable 5. It is the better everyday product decision. It replaces Opus 4.8, wins enough important tests to challenge the absolute frontier, and makes careful AI judgment affordable for ordinary Tuesday work—not only for the one project important enough to receive a blank cheque.

02

Strengths and honest limitations

Key Strengths

  • Knowledge work is the headline: Opus 5 reaches 1861 Elo on GDPval-AA v2 in Anthropic’s launch comparison, ahead of Fable 5 at 1747 and GPT-5.6 Sol at 1736. Early customers also report stronger finance, legal, spreadsheet, and due-diligence work.
  • It can operate, not only answer: The model scores 70.6% on OSWorld 2.0 and 26.0% on AutomationBench, leading the compared models. That is evidence for navigating software and completing business workflows, not merely composing a clever paragraph.
  • Novel problems are a real strength: Its 30.2% on ARC-AGI-3 is almost four times GPT-5.6 Sol’s 7.8% in Anthropic’s table. Artificial Analysis then placed max effort first on its broader Intelligence Index at 61.
  • Fable-like quality without the Fable tax: API pricing is $5/$25 per million input/output tokens—half of Fable 5. Opus 5 also has no general-access data-retention requirement, while Fable can require retention acceptance in some deployments.
  • A serious professional workspace: One million tokens of context, 128k output, strong document and chart vision, multi-sheet spreadsheet work, slide creation, web search, files, PDFs, memory, and computer use make it useful beyond chat.

Honest Limitations

  • This is a launch-week crown: Anthropic supplied most per-benchmark comparisons, and independent leaderboards are still filling in. Artificial Analysis’ #1 result is valuable, but one composite score is not a permanent constitution.
  • ChatGPT remains the larger consumer platform: Claude lacks native image generation and does not match ChatGPT’s combined chat, Work, Codex, Sites, plugin, browser, and desktop footprint. Opus 5 leads our quality/value score; GPT-5.6 may still be the easier single subscription.
  • Fable keeps specialist wins: Fable 5 is fractionally ahead on Humanity’s Last Exam without tools, FrontierCode, CursorBench, and the held-out Legal Agent Benchmark; Mythos remains stronger in offensive cyber and autonomous biology.
  • Maximum effort is hungry: Artificial Analysis measured max effort as very verbose, while xhigh was slower than average. Start with high or xhigh only when the task earns it, and compare total finished-task cost rather than token list price.
  • Some integrations need adjustment: API web fetch and Priority Tier are missing at launch. Thinking defaults on, sampling controls are restricted, and Claude usage caps can still interrupt a long consumer workflow.
03

Benchmark Snapshot

GDPval-AA v2 — 1861 Elo (#1)

Anthropic's knowledge-work comparison puts Opus 5 ahead of Fable 5 at 1747, GPT-5.6 Sol at 1736, and Opus 4.8 at 1593.

ARC-AGI-3 — 30.2% (#1)

Novel interactive problem-solving. Opus 5 scores nearly four times GPT-5.6 Sol's 7.8% and far above Opus 4.8's 1.5% in the launch table.

OSWorld 2.0 — 70.6% (#1)

Computer-use tasks across desktop applications. Opus 5 leads Fable 5 at 66.1% and GPT-5.6 Sol at 62.6% in Anthropic's comparison.

AutomationBench — 26.0% (#1)

End-to-end business workflows. The result is about 1.5× Fable 5's 17.4% at the published setup.

Artificial Analysis Intelligence Index — 61 (#1)

Independent max-effort composite across nine evaluations. High and xhigh score 59 and 60; max also used many tokens, so effort choice matters.

04

The Verdict

Claude Opus 5 takes our #1 everyday-intelligence score, narrowly ahead of GPT-5.6 and clearly above Fable 5. The reason is not that it wins every exam. It is that it leads the most relevant mix of knowledge work, novel problem-solving, browser search, computer use, and workflow completion, then charges half Fable’s price with fewer deployment restrictions. Choose ChatGPT if one enormous consumer ecosystem matters most; choose Fable only for a narrow peak where its slight lead repays double the price; choose Opus 5 when the work is ambiguous, document-heavy, and important enough to deserve judgment.

05

Frequently Asked Questions