Ranked #3 Coding — AI That Writes Production Code
Anthropic

Claude Fable 5.1

The Mythos-class reasoning engine refined for agentic coding. Same weights as the restricted Mythos 5.1, but optimized with a 75% cache discount that transforms long-horizon engineering economics. Best deployed in Claude Code, where its AA Coding Agent score of 70 leads the field.

ReasoningLong-ContextAgentic
9.8out of 10
Official Website
Best for

The Mythos-class reasoning engine refined for agentic coding. Same weights as the restricted Mythos 5.1, but optimized with a 75% cache discount that transforms long-horizon engineering economics. Best deployed in Claude Code, where its AA Coding Agent score of 70 leads the field.

Why It Wins

AA Coding Agent 70 (in Claude Code). CursorBench 3.2 73.4%. FrontierSWE v2 0.57. Terminal-Bench 4.0 55.8%. Cache reads slashed 75% to $0.25/MTok, unlocking deep repository searches.

Watch out

Max effort can degrade SWE performance through over-editing (medium effort beats max on FrontierCode). DeepSWE 1.1 score dipped to 67.4% due to over-implementation. High base cost ($10/$50) rapidly depletes limits on cold caches.

01

What It Actually Is

When Fable 5 launched, it proved that a reasoning-heavy Mythos-class architecture could dominate software engineering benchmarks if given enough time to think. Claude Fable 5.1 doesn’t reinvent that wheel; instead, it refines the economics and integration of that reasoning process. By combining the same weights as the highly restricted Mythos 5.1 with a massive 75% discount on cache reads, Anthropic has turned a prohibitively expensive long-horizon agent into a practical tool for deep repository work.

In the right harness, Fable 5.1 is virtually unmatched. Inside Claude Code, it posts an Artificial Analysis Coding Agent score of 70, establishing itself as the premier autonomous assistant. It pushes the state-of-the-art on CursorBench 3.2 to 73.4% and jumps to 55.8% on Terminal-Bench 4.0—a massive improvement over Fable 5’s 42.0%. On FrontierSWE v2, which measures unconstrained engineering execution, its 0.57 score comfortably leads Opus 5 (0.52) and leaves competitors like Astra Sol (0.32) far behind. Power users, particularly those navigating complex Jane Street-style codebases, report that Fable 5.1 excels at maintaining coherence across messy, 40-file execution traces without losing the plot.

However, Fable 5.1’s always-on adaptive thinking comes with a significant caveat: more effort can actually degrade software engineering performance. The system card data confirms that on FrontierCode, “medium” effort (50.9%) outperforms “max” effort (50.3%). This is because at max effort, Fable 5.1 tends to over-scope solutions—making drive-by file edits, adding unnecessary documentation, and engineering complex CI jobs when a simple bug fix was requested. This over-implementation is why its DeepSWE 1.1 score dipped slightly to 67.4%. To get the best out of Fable 5.1, you need to constrain it; pin it to medium effort for standard tasks, and only unleash max effort when you have robust tests already in place to guide its reasoning.

The operational costs also require careful management. While the $0.25/MTok cache read price is a revelation, Fable 5.1 is notoriously verbose, generating roughly 1.7× more tokens than its predecessor. MineBench users have reported seeing 3× the cost and double the latency compared to Fable 5. Furthermore, the base price remains $10/$50 per million tokens; dropping a massive repository into a cold cache will rapidly deplete your Claude usage limits. And because it runs the standard safety classifiers, tasks touching offensive security or dual-use development will silently route to Opus.

Claude Fable 5.1 is not the model you want auto-completing every line of your React component. Astra leads in computer-use and browser automation, and Opus 5 is significantly cheaper for daily, medium-effort coding. But when you face a deeply entrenched architectural bug, need a senior-level planner, or require a model that can synthesize a massive repository over a multi-hour session, Fable 5.1 is the tool you bring to the fight.

02

Strengths and honest limitations

Key Strengths

  • The Claude Code Champion: When harnessed in Anthropic’s own Claude Code, Fable 5.1 achieves an Artificial Analysis Coding Agent score of 70, outperforming both its predecessor and Codex-based alternatives.
  • CursorBench & Terminal Mastery: Scores 73.4% on CursorBench 3.2 and 55.8% on Terminal-Bench 4.0, cementing its position as a top-tier agent for deeply integrated IDE tasks.
  • Cache Economics for SWE: The 75% reduction in cache read pricing ($0.25/MTok) makes iterative debugging and long-context repository-wide searches economically viable for everyday engineering.
  • FrontierSWE Leader: Achieves a 0.57 on FrontierSWE v2, demonstrating superior capability in handling complex, unconstrained software engineering tasks compared to Opus 5 (0.52).
  • Jane Street-Style Endurance: Power users report that Fable 5.1 solves more complex coding problems than Opus 5 and, crucially, stays readable and coherent over long, multi-file execution traces.

Honest Limitations

  • The Over-Editing Trap: More effort doesn’t mean better code. FrontierCode testing shows ‘medium’ effort (50.9%) beats ‘max’ (50.3%) because the model tends to make drive-by edits and over-scope solutions.
  • Token Hunger & Latency: The model is notoriously verbose. MineBench users saw ~3× the cost and 2× the latency compared to Fable 5, as the model generates ~1.7× more tokens on average.
  • Astra Wins the OS/Browser Wars: While Fable 5.1 excels in raw coding, it falls short of Astra in OSWorld 2.0 and computer-use tasks, making it less ideal for end-to-end browser automation.
  • Cold Cache Caps: The $10/MTok cold read price means that dropping a massive repository in without warming the cache will rapidly burn through Claude usage caps and API budgets.
  • Safety Classifiers Intrude: It can find vulnerabilities but won’t write exploits. Legitimate offensive security research or dual-use coding tasks will be automatically routed to Opus.
03

Benchmark Snapshot

AA Coding Agent — 70

Achieved inside the Claude Code harness, establishing it as the premier agentic coding assistant.

CursorBench 3.2 — 73.4%

A solid improvement over Fable 5 (70.5%), proving its value in long-horizon IDE tasks.

FrontierSWE v2 — 0.57

Leads Opus 5 (0.52) and Astra Sol (0.32) in unconstrained software engineering challenges.

Terminal-Bench 4.0 — 55.8%

Massive jump from Fable 5's 42.0%, though trailing the unrestricted Mythos 5.1 (60.9%).

04

The Verdict

Claude Fable 5.1 is a scalpel that occasionally tries to be a chainsaw. On the benchmarks that matter for agentic coding — CursorBench, FrontierSWE, and AA Coding Agent — it is an undeniable heavyweight. The 75% cache read discount makes long-horizon repository work viable for the first time. But its tendency to over-implement at high effort, combined with significant token hunger, means it requires a disciplined prompt and harness (like Claude Code) to shine. Anthropic’s advice holds: use Opus 5 for your daily medium-effort work, but when you need a senior reviewer, a planner, or have a complex bug with tests already written, tag in Fable 5.1.