Most chatbots can feel like a brilliant intern on a frantic Monday: quick, eager, and perfectly capable of presenting a polished spreadsheet whose grand total is quietly wrong. Claude Opus 5 feels more like the colleague who checks the formulas before pressing Send.
That judgment is the reason it moves straight to #1 in our everyday category. Anthropic’s launch table gives it 1861 Elo on GDPval-AA v2, a knowledge-work evaluation built around useful professional tasks. Fable 5 scores 1747; GPT-5.6 Sol, 1736; Opus 4.8, 1593. Box reports an 8% overall improvement over 4.8, with larger gains in data analysis and due diligence. A financial-evaluation customer saw nine points more accuracy with fewer turns and 60% less time. These are early customer reports, but they describe the same shape as the benchmarks: careful work with fewer trips around the block.
More than a well-read typewriter
Opus 5 scores 70.6% on OSWorld 2.0, which asks a model to use a computer, not merely talk about one. It reaches 26.0% on AutomationBench, roughly one and a half times the next result in Anthropic’s comparison. Zapier reports that Opus 5 completed an entire customer-retention workflow: it read account-health data, identified customers at risk of leaving, alerted the people responsible for those accounts, and prepared a summary. Earlier models failed to finish that chain.
The strangest number is 30.2% on ARC-AGI-3. This is an interactive test of solving unfamiliar problems, closer to learning the rules of a new board game than recalling a fact. GPT-5.6 Sol scored 7.8% in the same table. A benchmark is not a soul, but this one suggests Opus 5 is unusually good when the instructions stop holding its hand.
Independent testing supports the direction. Artificial Analysis gives Opus 5 max effort 61, currently first on its Intelligence Index. High scores 59 and xhigh 60. The ladder matters: you can buy more thought when a contract or financial model deserves it, then turn the dial down for ordinary drafting.
Why it ranks above Fable 5
Fable is still slightly better in a few narrow places: Humanity’s Last Exam without tools, FrontierCode, the highest CursorBench setting, and a held-out legal-agent test. If your private evaluation matches one of those lanes, use it. But Fable costs $10/$50 per million input/output tokens. Opus 5 costs $5/$25, has no general-access data-retention requirement, and its safety classifiers are expected to intervene much less often.
Think of Fable as a racing prototype and Opus 5 as the road car that is almost as fast, costs half as much to run, and can legally travel far beyond the track. For daily work, the road car is the more useful machine.
Why ChatGPT still belongs in the conversation
GPT-5.6 remains a superb all-round system. ChatGPT combines Work, Codex, image generation, Sites, connected apps, a browser, and desktop computer use in one enormous ecosystem. Claude cannot generate images natively, and its plan limits can be tighter. Our ranking says Opus 5 offers the best current blend of reasoning quality, professional agency, price, and deployability; it does not say every person should cancel ChatGPT.
The hardware behind that workbench is substantial: 1M tokens of context by default, 128k maximum output, vision for charts and documents, PDFs, files, memory, web search, tools, and computer use. The API model is claude-opus-5, with support across Anthropic’s platform and major clouds. Fast mode runs around 2.5 times faster at twice the base token price.
There are caveats. Independent data is still young. Max effort can be verbose, and xhigh has noticeable first-token latency. API web fetch and Priority Tier are absent at launch. The sensible habit is to start lower, verify the result, and raise effort only when the cost of being wrong is larger than the cost of thinking.
Opus 5 is not merely a cheaper Fable 5. It is the better everyday product decision. It replaces Opus 4.8, wins enough important tests to challenge the absolute frontier, and makes careful AI judgment affordable for ordinary Tuesday work—not only for the one project important enough to receive a blank cheque.