Claude Fable 5.1 Comprehensive Analysis: Benchmarks, Improvements, Pricing, and Speed

Claude Fable 5.1's complete results across 7 mainstream benchmark categories, with item-by-item comparisons against Fable 5, Opus 5, and GPT-5.6 Sol, along with explanations of 5.1's long-task performance, tool calling, cache costs, speed, plan limits, and migration changes.

PandaNpcFirst published on
Claude Fable 5.1 Comprehensive Analysis: Benchmarks, Improvements, Pricing, and Speed

Claude Fable 5.1 was officially released on September 1, 2026. The bottom line: this is not a routine model upgrade for ordinary Q&A — it is a high-cost flagship model built for long-horizon coding, complex research, cross-application tool use, and unattended agents. Across the seven mainstream benchmark categories Anthropic published, Fable 5.1 beats Fable 5 across the board; however, its API unit price is still twice Opus 5's, and the official relative latency is labeled "Slower."

If your task is just editing a file, writing a short passage, or answering everyday questions, Anthropic still recommends starting with Opus 5. Fable 5.1 is better suited to long-chain tasks that ordinary models cannot reliably complete even at high effort. This article puts the official release data, the meaning of each benchmark, 5.1's specific improvements, pricing, speed, and migration constraints together, so you don't just look at a single "who wins" benchmark chart.

Fable 5.1 Specifications at a Glance

Item Claude Fable 5.1 Claude Opus 5 Claude Sonnet 5
Context window 1M token 1M token 1M token
Maximum output 128K token 128K token 128K token
Input price $10 / MTok $5 / MTok $2 / MTok
Output price $50 / MTok $25 / MTok $10 / MTok
Relative latency Slower Moderate Fast
Thinking Adaptive, always on Adaptive Adaptive
Default effort High High High
Reliable knowledge cutoff June 2026 May 2026 January 2026

The "Slower" label here comes from the Anthropic model catalog. It denotes a relative latency tier, not a fixed tokens-per-second number. How long a task actually takes depends on effort, the number of tool-call rounds, cache hit rate, context length, and failed retry count.

Fable 5.1 long-horizon agent workflow from planning to verification
Fable 5.1 is not aimed at one-shot answers, but at long tasks that continuously plan, execute, recover, verify, and report.

Fable 5.1 Mainstream Benchmark Results, Full Table

The table below is taken directly from the main table on Anthropic's September 1, 2026 release page. To avoid misreading, percentages, the GDPval-AA raw score, OSWorld's partial/strict results, and HLE's with-tools/no-tools results are all listed separately. Each benchmark keeps its official English name, with the corresponding translated name on the next line; localized versions also display their language's name below the English name.

Capability Benchmark Fable 5.1 Fable 5 Opus 5 GPT-5.6 Sol 5.1 vs. Fable 5
Agentic scientific research Terminal-Bench-Science 0.1
Agentic scientific research terminal benchmark
52.6% 24.7% 29.0% 22.4% +27.9 percentage points
Agentic terminal coding Terminal-Bench 4.0
Agentic terminal coding benchmark
55.8% 42.0% 52.3% 37.3% +13.8 percentage points
Professional knowledge work GDPval-AA v2
Real-world economically valuable professional task benchmark
1853 1723 1824 1711 +130 points
Computer operation OSWorld 2.0 — Partial
Real computer operation benchmark: partial completion score
77.9% 72.9% 75.4% — +5.0 percentage points
Computer operation OSWorld 2.0 — Strict
Real computer operation benchmark: strict completion rate
41.7% 36.1% 39.6% — +5.6 percentage points
Multidisciplinary reasoning Humanity's Last Exam — No tools
Humanity's Last Exam: no tools
60.9% 57.8% 56.6% — +3.1 percentage points
Multidisciplinary reasoning Humanity's Last Exam — With tools
Humanity's Last Exam: with tools
65.0% 63.8% 63.6% — +1.2 percentage points
Business workflows AutomationBench
Cross-application business automation workflow benchmark
31.4% 17.1% 26.9% 19.6% +14.3 percentage points
Agentic coding CursorBench 3.2.0
Real multi-file coding task benchmark
73.4% 70.5% 70.0% 67.2% +2.9 percentage points

— means Anthropic's main table does not provide a score for that entry; it does not mean the model cannot complete the task. On Terminal-Bench 4.0, Mythos 5.1, which has fewer restrictions and is open only to trusted programs, scored 60.9%. This article focuses on Fable 5.1, which is available to ordinary users, and does not mix Mythos's results into the Fable column.

What Terminal-Bench-Science 0.1 Measures

Terminal-Bench-Science 0.1 contains 70 real research workflows across five domains: life sciences, physical sciences, earth sciences, mathematics, and engineering. The tasks are not multiple-choice; agents must perform data analysis, statistical inference, simulation, optimization, theorem proving, image reconstruction, or scientific machine learning in a terminal, and reproducible tests check the outputs.

Fable 5.1's 52.6%, versus Fable 5's 24.7% in the official reproduction runs, is a 27.9-percentage-point improvement — the most dramatic jump in the whole table. But the official note also gives a standard error of roughly ±3.5–4.5 percentage points per model, so the first decimal place should not be treated as an absolutely precise ranking.

What Terminal-Bench 4.0 and CursorBench 3.2.0 Measure

Terminal-Bench has agents handle complex tasks in real terminal environments, focusing on whether the agent understands the environment, calls tools, modifies files, and ultimately passes verification. Fable 5.1 scores 55.8%, which is 13.8 percentage points above Fable 5 and also above Opus 5's 52.3%.

CursorBench 3.2 draws from ambiguous, multi-file tasks in real Cursor sessions; version 3.2 adds instruction-following and advanced tool-use questions. Fable 5.1 Max scores 73.4%, and while the margin is less dramatic than on Terminal-Bench-Science, it averages 72,060 tokens and $9.64 per task, versus Fable 5 Max's 103,525 tokens and $17.32. What is more worth noting here is the drop in tokens and cost when completing comparable tasks, rather than the 2.9 percentage points by themselves.

What GDPval-AA v2 Measures

GDPval-AA v2 uses 220 tasks designed by industry professionals across 44 occupations and 9 industries to assess how well a model handles real, economically valuable knowledge work. Its 1853 is a composite score, not 1853%. This result indicates that Fable 5.1 beats Fable 5 — and narrowly beats Opus 5 — on professional documents, analysis, and deliverables.

Why OSWorld 2.0 Has Two Scores

OSWorld 2.0 includes 108 long-horizon computer operation workflows, and the median human completion time per task is about 1.6 hours. Partial gives partial credit according to multiple checkpoints; Strict counts a task only when the goal is fully achieved.

So Fable 5.1's 77.9% partial and 41.7% strict are not contradictory: it often completes most steps, yet still leaves some tasks unfinished end to end. It also reminds us that a computer agent that "looks like it is constantly working" is not the same as a task having actually succeeded.

What Humanity's Last Exam Measures

Humanity's Last Exam, built by the Center for AI Safety and Scale AI, contains 2,500 cross-disciplinary, expert-level questions that are difficult to answer with simple retrieval. Fable 5.1 beats Fable 5 by 3.1 percentage points without tools, but by only 1.2 percentage points with tools. It is genuinely stronger — but not every benchmark shows double-digit growth.

What AutomationBench Measures

AutomationBench checks whether an agent can complete sales, marketing, operations, customer support, finance, and HR workflows across simulated SaaS tools such as CRM, email, calendar, messaging, and project management. Fable 5.1 improves from 17.1% to 31.4%, a relative gain of about 84%, showing real progress in multi-application state management and long-step execution. Still, 31.4% also shows that unattended business automation of this kind remains far from solved.

Fable 5.1 capability coverage matrix across seven mainstream benchmarks
The seven benchmarks cover scientific research, terminal coding, professional knowledge work, computer operation, multidisciplinary reasoning, business workflows, and multi-file coding.

What Has Improved in Fable 5.1 Compared with Fable 5

1. Less Likely to Cut Corners on Long Tasks

Anthropic positions Fable 5.1 as a long-horizon agent model: plan first, then use tools, recover after failure, verify results, and proactively report progress. It emphasizes fixing root causes rather than making local workarounds just to turn one test green as quickly as possible. Official target scenarios include cross-codebase features, code review, performance optimization, multi-day autonomous sessions, research, documents, spreadsheets, and slide decks.

2. Low Effort Can Approach the Previous Generation's High-Cost Results

Fable 5.1 adds the ability to adjust effort per message. The official cost curves show Low or Medium effort matching or exceeding Fable 5 on some tasks while using fewer tokens and lower cost. Claude Code defaults to High effort; Claude.ai and Cowork default to Medium. When comparing models, you must confirm the effort setting first, and results from different levels cannot be directly compared.

3. Readable Progress Between Tool Calls

The API adds an experimental display: "updates" capability that lets the model give user-facing progress updates between tool calls. For tasks that run for hours, this makes it easier to tell what the agent is doing and whether it has gone off track, compared with only seeing tool cards appear continuously.

4. Fuller Writing, Visual Verification, and Self-Testing

According to the official announcement, 5.1 proactively writes tests to verify its own changes and uses vision capabilities to check output against design goals. Partner feedback also centers on cleaner progress updates, more complete answers to multi-part questions, tracing call chains across multiple services, and staying readable after long-running sessions. These are qualitative observations and should not be passed off as unified benchmark conclusions.

5. Safety Blocks Are More Precise, but Not Gone

Fable 5.1's cybersecurity guardrails reduce false positives by about 60% compared with Fable 5's launch, and downgrade triggers for basic biology and medicine questions drop by around 85%. It can now be used to discover software vulnerabilities, but dual-use tasks such as penetration testing, exploit-code generation, and binary vulnerability scanning may still be routed to Opus. Routes are not billed at Fable prices.

6. Content Provenance Marking Added

Fable 5.1 introduces content provenance. Generated text may carry a statistical watermark that can be verified through Anthropic's content-checking tool. It will not place a visible marker in the text itself, but for content platforms, education, and enterprise audits, this is a new behavior worth understanding in advance.

Six core improvements in Fable 5.1 over the prior generation
The changes in 5.1 center on long tasks, low-effort efficiency, progress reporting, self-verification, fewer false positives, and content provenance.

How Fable 5.1 Costs Are Billed

Billing item Fable 5.1 price Compared with Fable 5
Input $10 / 1M tokens Same
Output $50 / 1M tokens Same
5-minute cache write $12.50 / 1M tokens Same tier
1-hour cache write $20 / 1M tokens Same tier
Cache read $0.25 / 1M tokens 75% lower
Batch API 50% of the input/output prices Per Batch API rules

Suppose one long task reads 10M cached tokens in total, generates 200K new input tokens and 50K output tokens, and we ignore cache writes:

text
缓存读取:10 × $0.25 = $2.50
新输入:0.2 × $10 = $2.00
输出:0.05 × $50 = $2.50
合计:$7.00

The same 10M cache reads on Fable 5 would be about $10, so this line item alone differs by $7.50. Based on this, Anthropic estimates that a typical token-billed task can see its real cost drop by about 25%, and up to roughly 45% for highly agentic tasks that reuse long contexts frequently. This is not a broad cut to the API list price — it is the more cache hits, the larger the savings.

Subscribers Do Not Necessarily Get the API Cache Discount Directly

Max and Team/Enterprise premium seats can use up to 50% of their weekly quota on Fable models; Pro and standard seats start using usage credits from the beginning. The exact usage shown on the Usage page of your account governs this. The cache-read price cut is mainly a change for token-billed scenarios, so you cannot directly convert the API's 75% cache discount into "four times more usage on Max."

Fable 5.1 input, output, and cache cost structure
Fable 5.1's input and output unit prices are unchanged; the main reduction comes from cache reads, and the more long-context reuse, the greater the benefit.

Is Fable 5.1 Actually Fast?

The answer: single-turn responses are slower, but total completion time for complex tasks can be shorter.

  • Anthropic's model catalog marks Fable 5.1 as Slower, Opus 5 as Moderate, and Sonnet 5 as Fast.
  • Claude Code defaults to High effort, which naturally uses more reasoning time than Low or Medium.
  • Every reports that, on its own workloads, Fable 5.1 is about twice as fast as Opus 5 with half the tokens; this is a partner-specific test, not a general speed guarantee.
  • On CursorBench, Fable 5.1 Max averages 70 steps and 72,060 tokens, while Fable 5 Max also takes 72 steps but uses 103,525 tokens. 5.1's advantage reads more like "fewer detours and fewer tokens" than a faster first token.

For chat, short code, and highly interactive work, choose Sonnet 5 or Opus 5 first; for cross-codebase debugging, long-horizon research, and unattended tasks that need self-verification, consider Fable 5.1.

Compatibility Changes to Handle Before Upgrading to Fable 5.1

Upgrade Claude Code to at Least 2.1.250

Anthropic's plan and availability notes require Claude Code 2.1.250 or later. If you cannot see Fable 5.1, first check:

bash
claude --version

Then update using the official Claude Code upgrade method and restart the session. Fable 5.1's API model ID is:

text
claude-fable-5-1

Forced Tool Use May Return Errors Directly

If a request forces the model to call a specific tool, Fable 5.1 may return an error. Before migrating, audit tool_choice and any forced-tool flows in your own agents; do not assume Fable 5's behavior is fully compatible.

Thinking Blocks Cannot Be Reused Arbitrarily Across Models

Earlier models cannot read Fable 5.1's thinking blocks. When switching models, let the SDK handle these blocks according to the official rules rather than passing them verbatim to another model.

Editing Old History Can Return a 400

For new API accounts created after August 31, 2026, thinking blocks are only valid under the original conversation prefix where they were generated. Replaying thinking blocks after editing system, tools, or earlier messages can return a 400. Anthropic recommends keeping history append-only; when compaction is necessary, replace the entire old history with a new summary instead of editing old turns while keeping the original thinking blocks. For the detailed pattern, see the Fable 5.1 Prompting and Migration Guide.

When Should You Choose Fable 5.1

Good fits:

  • Long tasks across multiple repositories, services, or applications;
  • Agents that must run for hours and recover from failures by themselves;
  • Scientific research, complex data analysis, and multi-stage professional deliverables;
  • Difficult-to-reproduce system problems, performance optimization, and root-cause investigation;
  • API workloads that heavily reuse long contexts and have a high cache-read share.

Poor fits:

  • Simple chat, short writing, or small single-file edits;
  • Real-time interaction that is highly sensitive to time to first token;
  • High-output tasks with a fixed budget but no cache reuse;
  • Scenarios that cannot accept the default 30-day security-monitoring data retention and do not qualify for the enterprise exception;
  • Workflows that expect every cybersecurity or life-sciences request to avoid downgrades.

How to View Long Tasks Remotely

Fable 5.1's value often appears in tasks that run for hours or even days, but that does not mean someone must stay in front of the development machine. First confirm that Fable 5.1 is available in local Claude Code 2.1.250+, then use PandaNpc Agent to view Claude Code sessions on the same development machine from a browser, desktop, or phone, handle interactions, and continue sending commands.

What changes is the control entry point; the code, tools, and builds still run on your development machine. You can read the Claude Code Remote Access Guide and the Connecting to Claude Code from a Phone tutorial first, then hand long tasks to Fable 5.1 once the remote link is confirmed.

FAQ

How Much Stronger Is Fable 5.1 Than Fable 5?

It varies greatly by task. It improves by 27.9 percentage points on Terminal-Bench-Science and 14.3 percentage points on AutomationBench; on HLE with tools, it improves by only 1.2 percentage points. You cannot use a single largest gain to represent all tasks.

Is Fable 5.1 Faster Than Opus 5?

In the official relative-latency table, Fable 5.1 is "Slower" and Opus 5 is "Moderate." Some partners observe that Fable 5.1 is faster on complete tasks because it uses fewer tokens and less rework. Those two conclusions are not measuring the same thing.

Is Fable 5.1 Better Suited to Coding Than GPT-5.6 Sol?

In Anthropic's release table, Fable 5.1 is higher on Terminal-Bench-Science, Terminal-Bench, AutomationBench, and CursorBench; but different models use different agent harnesses, effort levels, tools, and prices, so this is not enough to say Fable wins on every real repository. You can refer to the GPT-5.6 Sol 1M Context Configuration Guide and then A/B test with your own representative tasks.

Why Does Fable 5.1 Suddenly Switch to Opus?

Some cybersecurity and life-sciences requests are routed to Opus by safeguards. A model switch is not necessarily a fault; Anthropic states that routed requests are not billed at Fable prices.

Does Fable 5.1 Have a Watermark?

It has a statistical content-provenance mechanism, but it does not add a visible label to the surface of the text. It is not the same as a visible watermark in the corner of an image.

Summary

Claude Fable 5.1's clearest gains are concentrated in scientific research, terminal tasks, and cross-application business workflows: it is better at sustained execution, failure recovery, and result verification, and it lowers long-task costs through cheaper cache reads. But it is still expensive, has slower single-turn latency, and brings compatibility changes around thinking blocks, forced tool use, and history editing.

The right question is not "Is it number one on the benchmarks?" but rather: is your task long and complex enough, and are the costs of failure and rework high enough, to justify Fable 5.1? If the answer is no, starting from Opus 5 or Sonnet 5 is usually more appropriate.