Claude Fable 5.1 Comprehensive Analysis: Benchmarks, Improvements, Pricing, and Speed
Claude Fable 5.1's complete results across 7 mainstream benchmark categories, with item-by-item comparisons against Fable 5, Opus 5, and GPT-5.6 Sol, along with explanations of 5.1's long-task performance, tool calling, cache costs, speed, plan limits, and migration changes.

Claude Fable 5.1 was officially released on September 1, 2026. The bottom line: this is not a routine model upgrade for ordinary Q&A — it is a high-cost flagship model built for long-horizon coding, complex research, cross-application tool use, and unattended agents. Across the seven mainstream benchmark categories Anthropic published, Fable 5.1 beats Fable 5 across the board; however, its API unit price is still twice Opus 5's, and the official relative latency is labeled "Slower."
If your task is just editing a file, writing a short passage, or answering everyday questions, Anthropic still recommends starting with Opus 5. Fable 5.1 is better suited to long-chain tasks that ordinary models cannot reliably complete even at high effort. This article puts the official release data, the meaning of each benchmark, 5.1's specific improvements, pricing, speed, and migration constraints together, so you don't just look at a single "who wins" benchmark chart.
Fable 5.1 Specifications at a Glance
| Item | Claude Fable 5.1 | Claude Opus 5 | Claude Sonnet 5 |
|---|---|---|---|
| Context window | 1M token | 1M token | 1M token |
| Maximum output | 128K token | 128K token | 128K token |
| Input price | $10 / MTok | $5 / MTok | $2 / MTok |
| Output price | $50 / MTok | $25 / MTok | $10 / MTok |
| Relative latency | Slower | Moderate | Fast |
| Thinking | Adaptive, always on | Adaptive | Adaptive |
| Default effort | High | High | High |
| Reliable knowledge cutoff | June 2026 | May 2026 | January 2026 |
The "Slower" label here comes from the Anthropic model catalog. It denotes a relative latency tier, not a fixed tokens-per-second number. How long a task actually takes depends on effort, the number of tool-call rounds, cache hit rate, context length, and failed retry count.
Fable 5.1 Mainstream Benchmark Results, Full Table
The table below is taken directly from the main table on Anthropic's September 1, 2026 release page. To avoid misreading, percentages, the GDPval-AA raw score, OSWorld's partial/strict results, and HLE's with-tools/no-tools results are all listed separately. Each benchmark keeps its official English name, with the corresponding translated name on the next line; localized versions also display their language's name below the English name.
| Capability | Benchmark | Fable 5.1 | Fable 5 | Opus 5 | GPT-5.6 Sol | 5.1 vs. Fable 5 |
|---|---|---|---|---|---|---|
| Agentic scientific research | Terminal-Bench-Science 0.1 Agentic scientific research terminal benchmark |
52.6% | 24.7% | 29.0% | 22.4% | +27.9 percentage points |
| Agentic terminal coding | Terminal-Bench 4.0 Agentic terminal coding benchmark |
55.8% | 42.0% | 52.3% | 37.3% | +13.8 percentage points |
| Professional knowledge work | GDPval-AA v2 Real-world economically valuable professional task benchmark |
1853 | 1723 | 1824 | 1711 | +130 points |
| Computer operation | OSWorld 2.0 — Partial Real computer operation benchmark: partial completion score |
77.9% | 72.9% | 75.4% | — | +5.0 percentage points |
| Computer operation | OSWorld 2.0 — Strict Real computer operation benchmark: strict completion rate |
41.7% | 36.1% | 39.6% | — | +5.6 percentage points |
| Multidisciplinary reasoning | Humanity's Last Exam — No tools Humanity's Last Exam: no tools |
60.9% | 57.8% | 56.6% | — | +3.1 percentage points |
| Multidisciplinary reasoning | Humanity's Last Exam — With tools Humanity's Last Exam: with tools |
65.0% | 63.8% | 63.6% | — | +1.2 percentage points |
| Business workflows | AutomationBench Cross-application business automation workflow benchmark |
31.4% | 17.1% | 26.9% | 19.6% | +14.3 percentage points |
| Agentic coding | CursorBench 3.2.0 Real multi-file coding task benchmark |
73.4% | 70.5% | 70.0% | 67.2% | +2.9 percentage points |
— means Anthropic's main table does not provide a score for that entry; it does not mean the model cannot complete the task. On Terminal-Bench 4.0, Mythos 5.1, which has fewer restrictions and is open only to trusted programs, scored 60.9%. This article focuses on Fable 5.1, which is available to ordinary users, and does not mix Mythos's results into the Fable column.
What Terminal-Bench-Science 0.1 Measures
Terminal-Bench-Science 0.1 contains 70 real research workflows across five domains: life sciences, physical sciences, earth sciences, mathematics, and engineering. The tasks are not multiple-choice; agents must perform data analysis, statistical inference, simulation, optimization, theorem proving, image reconstruction, or scientific machine learning in a terminal, and reproducible tests check the outputs.
Fable 5.1's 52.6%, versus Fable 5's 24.7% in the official reproduction runs, is a 27.9-percentage-point improvement — the most dramatic jump in the whole table. But the official note also gives a standard error of roughly ±3.5–4.5 percentage points per model, so the first decimal place should not be treated as an absolutely precise ranking.
What Terminal-Bench 4.0 and CursorBench 3.2.0 Measure
Terminal-Bench has agents handle complex tasks in real terminal environments, focusing on whether the agent understands the environment, calls tools, modifies files, and ultimately passes verification. Fable 5.1 scores 55.8%, which is 13.8 percentage points above Fable 5 and also above Opus 5's 52.3%.
CursorBench 3.2 draws from ambiguous, multi-file tasks in real Cursor sessions; version 3.2 adds instruction-following and advanced tool-use questions. Fable 5.1 Max scores 73.4%, and while the margin is less dramatic than on Terminal-Bench-Science, it averages 72,060 tokens and $9.64 per task, versus Fable 5 Max's 103,525 tokens and $17.32. What is more worth noting here is the drop in tokens and cost when completing comparable tasks, rather than the 2.9 percentage points by themselves.
What GDPval-AA v2 Measures
GDPval-AA v2 uses 220 tasks designed by industry professionals across 44 occupations and 9 industries to assess how well a model handles real, economically valuable knowledge work. Its 1853 is a composite score, not 1853%. This result indicates that Fable 5.1 beats Fable 5 — and narrowly beats Opus 5 — on professional documents, analysis, and deliverables.
Why OSWorld 2.0 Has Two Scores
OSWorld 2.0 includes 108 long-horizon computer operation workflows, and the median human completion time per task is about 1.6 hours. Partial gives partial credit according to multiple checkpoints; Strict counts a task only when the goal is fully achieved.
So Fable 5.1's 77.9% partial and 41.7% strict are not contradictory: it often completes most steps, yet still leaves some tasks unfinished end to end. It also reminds us that a computer agent that "looks like it is constantly working" is not the same as a task having actually succeeded.
What Humanity's Last Exam Measures
Humanity's Last Exam, built by the Center for AI Safety and Scale AI, contains 2,500 cross-disciplinary, expert-level questions that are difficult to answer with simple retrieval. Fable 5.1 beats Fable 5 by 3.1 percentage points without tools, but by only 1.2 percentage points with tools. It is genuinely stronger — but not every benchmark shows double-digit growth.
What AutomationBench Measures
AutomationBench checks whether an agent can complete sales, marketing, operations, customer support, finance, and HR workflows across simulated SaaS tools such as CRM, email, calendar, messaging, and project management. Fable 5.1 improves from 17.1% to 31.4%, a relative gain of about 84%, showing real progress in multi-application state management and long-step execution. Still, 31.4% also shows that unattended business automation of this kind remains far from solved.
What Has Improved in Fable 5.1 Compared with Fable 5
1. Less Likely to Cut Corners on Long Tasks
Anthropic positions Fable 5.1 as a long-horizon agent model: plan first, then use tools, recover after failure, verify results, and proactively report progress. It emphasizes fixing root causes rather than making local workarounds just to turn one test green as quickly as possible. Official target scenarios include cross-codebase features, code review, performance optimization, multi-day autonomous sessions, research, documents, spreadsheets, and slide decks.
2. Low Effort Can Approach the Previous Generation's High-Cost Results
Fable 5.1 adds the ability to adjust effort per message. The official cost curves show Low or Medium effort matching or exceeding Fable 5 on some tasks while using fewer tokens and lower cost. Claude Code defaults to High effort; Claude.ai and Cowork default to Medium. When comparing models, you must confirm the effort setting first, and results from different levels cannot be directly compared.
3. Readable Progress Between Tool Calls
The API adds an experimental display: "updates" capability that lets the model give user-facing progress updates between tool calls. For tasks that run for hours, this makes it easier to tell what the agent is doing and whether it has gone off track, compared with only seeing tool cards appear continuously.
4. Fuller Writing, Visual Verification, and Self-Testing
According to the official announcement, 5.1 proactively writes tests to verify its own changes and uses vision capabilities to check output against design goals. Partner feedback also centers on cleaner progress updates, more complete answers to multi-part questions, tracing call chains across multiple services, and staying readable after long-running sessions. These are qualitative observations and should not be passed off as unified benchmark conclusions.
5. Safety Blocks Are More Precise, but Not Gone
Fable 5.1's cybersecurity guardrails reduce false positives by about 60% compared with Fable 5's launch, and downgrade triggers for basic biology and medicine questions drop by around 85%. It can now be used to discover software vulnerabilities, but dual-use tasks such as penetration testing, exploit-code generation, and binary vulnerability scanning may still be routed to Opus. Routes are not billed at Fable prices.
6. Content Provenance Marking Added
Fable 5.1 introduces content provenance. Generated text may carry a statistical watermark that can be verified through Anthropic's content-checking tool. It will not place a visible marker in the text itself, but for content platforms, education, and enterprise audits, this is a new behavior worth understanding in advance.
How Fable 5.1 Costs Are Billed
| Billing item | Fable 5.1 price | Compared with Fable 5 |
|---|---|---|
| Input | $10 / 1M tokens | Same |
| Output | $50 / 1M tokens | Same |
| 5-minute cache write | $12.50 / 1M tokens | Same tier |
| 1-hour cache write | $20 / 1M tokens | Same tier |
| Cache read | $0.25 / 1M tokens | 75% lower |
| Batch API | 50% of the input/output prices | Per Batch API rules |
Suppose one long task reads 10M cached tokens in total, generates 200K new input tokens and 50K output tokens, and we ignore cache writes:
缓存读取:10 × $0.25 = $2.50
新输入:0.2 × $10 = $2.00
输出:0.05 × $50 = $2.50
合计:$7.00The same 10M cache reads on Fable 5 would be about $10, so this line item alone differs by $7.50. Based on this, Anthropic estimates that a typical token-billed task can see its real cost drop by about 25%, and up to roughly 45% for highly agentic tasks that reuse long contexts frequently. This is not a broad cut to the API list price — it is the more cache hits, the larger the savings.
Subscribers Do Not Necessarily Get the API Cache Discount Directly
Max and Team/Enterprise premium seats can use up to 50% of their weekly quota on Fable models; Pro and standard seats start using usage credits from the beginning. The exact usage shown on the Usage page of your account governs this. The cache-read price cut is mainly a change for token-billed scenarios, so you cannot directly convert the API's 75% cache discount into "four times more usage on Max."
Is Fable 5.1 Actually Fast?
The answer: single-turn responses are slower, but total completion time for complex tasks can be shorter.
- Anthropic's model catalog marks Fable 5.1 as
Slower, Opus 5 asModerate, and Sonnet 5 asFast. - Claude Code defaults to High effort, which naturally uses more reasoning time than Low or Medium.
- Every reports that, on its own workloads, Fable 5.1 is about twice as fast as Opus 5 with half the tokens; this is a partner-specific test, not a general speed guarantee.
- On CursorBench, Fable 5.1 Max averages 70 steps and 72,060 tokens, while Fable 5 Max also takes 72 steps but uses 103,525 tokens. 5.1's advantage reads more like "fewer detours and fewer tokens" than a faster first token.
For chat, short code, and highly interactive work, choose Sonnet 5 or Opus 5 first; for cross-codebase debugging, long-horizon research, and unattended tasks that need self-verification, consider Fable 5.1.
Compatibility Changes to Handle Before Upgrading to Fable 5.1
Upgrade Claude Code to at Least 2.1.250
Anthropic's plan and availability notes require Claude Code 2.1.250 or later. If you cannot see Fable 5.1, first check:
claude --versionThen update using the official Claude Code upgrade method and restart the session. Fable 5.1's API model ID is:
claude-fable-5-1Forced Tool Use May Return Errors Directly
If a request forces the model to call a specific tool, Fable 5.1 may return an error. Before migrating, audit tool_choice and any forced-tool flows in your own agents; do not assume Fable 5's behavior is fully compatible.
Thinking Blocks Cannot Be Reused Arbitrarily Across Models
Earlier models cannot read Fable 5.1's thinking blocks. When switching models, let the SDK handle these blocks according to the official rules rather than passing them verbatim to another model.
Editing Old History Can Return a 400
For new API accounts created after August 31, 2026, thinking blocks are only valid under the original conversation prefix where they were generated. Replaying thinking blocks after editing system, tools, or earlier messages can return a 400. Anthropic recommends keeping history append-only; when compaction is necessary, replace the entire old history with a new summary instead of editing old turns while keeping the original thinking blocks. For the detailed pattern, see the Fable 5.1 Prompting and Migration Guide.
When Should You Choose Fable 5.1
Good fits:
- Long tasks across multiple repositories, services, or applications;
- Agents that must run for hours and recover from failures by themselves;
- Scientific research, complex data analysis, and multi-stage professional deliverables;
- Difficult-to-reproduce system problems, performance optimization, and root-cause investigation;
- API workloads that heavily reuse long contexts and have a high cache-read share.
Poor fits:
- Simple chat, short writing, or small single-file edits;
- Real-time interaction that is highly sensitive to time to first token;
- High-output tasks with a fixed budget but no cache reuse;
- Scenarios that cannot accept the default 30-day security-monitoring data retention and do not qualify for the enterprise exception;
- Workflows that expect every cybersecurity or life-sciences request to avoid downgrades.
How to View Long Tasks Remotely
Fable 5.1's value often appears in tasks that run for hours or even days, but that does not mean someone must stay in front of the development machine. First confirm that Fable 5.1 is available in local Claude Code 2.1.250+, then use PandaNpc Agent to view Claude Code sessions on the same development machine from a browser, desktop, or phone, handle interactions, and continue sending commands.
What changes is the control entry point; the code, tools, and builds still run on your development machine. You can read the Claude Code Remote Access Guide and the Connecting to Claude Code from a Phone tutorial first, then hand long tasks to Fable 5.1 once the remote link is confirmed.
FAQ
How Much Stronger Is Fable 5.1 Than Fable 5?
It varies greatly by task. It improves by 27.9 percentage points on Terminal-Bench-Science and 14.3 percentage points on AutomationBench; on HLE with tools, it improves by only 1.2 percentage points. You cannot use a single largest gain to represent all tasks.
Is Fable 5.1 Faster Than Opus 5?
In the official relative-latency table, Fable 5.1 is "Slower" and Opus 5 is "Moderate." Some partners observe that Fable 5.1 is faster on complete tasks because it uses fewer tokens and less rework. Those two conclusions are not measuring the same thing.
Is Fable 5.1 Better Suited to Coding Than GPT-5.6 Sol?
In Anthropic's release table, Fable 5.1 is higher on Terminal-Bench-Science, Terminal-Bench, AutomationBench, and CursorBench; but different models use different agent harnesses, effort levels, tools, and prices, so this is not enough to say Fable wins on every real repository. You can refer to the GPT-5.6 Sol 1M Context Configuration Guide and then A/B test with your own representative tasks.
Why Does Fable 5.1 Suddenly Switch to Opus?
Some cybersecurity and life-sciences requests are routed to Opus by safeguards. A model switch is not necessarily a fault; Anthropic states that routed requests are not billed at Fable prices.
Does Fable 5.1 Have a Watermark?
It has a statistical content-provenance mechanism, but it does not add a visible label to the surface of the text. It is not the same as a visible watermark in the corner of an image.
Summary
Claude Fable 5.1's clearest gains are concentrated in scientific research, terminal tasks, and cross-application business workflows: it is better at sustained execution, failure recovery, and result verification, and it lowers long-task costs through cheaper cache reads. But it is still expensive, has slower single-turn latency, and brings compatibility changes around thinking blocks, forced tool use, and history editing.
The right question is not "Is it number one on the benchmarks?" but rather: is your task long and complex enough, and are the costs of failure and rework high enough, to justify Fable 5.1? If the answer is no, starting from Opus 5 or Sonnet 5 is usually more appropriate.
Related guides

Codex Enable GPT-5.6 Sol 1M Context: 3-Line `config.toml` Configuration Tutorial
GPT-5.6 Sol officially supports a 1.05 million token context. Simply add 3 lines of configuration at the top of Codex's config.toml to run with a 1 million window, and it will automatically compact history at roughly 900,000 tokens. Includes permanent configuration, single command, verification, and cost reminders.
Read article →
How to Use Claude Code Remote Control from Your Phone (2026)
Use Claude Code Remote Control from your phone or browser. Learn the three official commands, setup limits, troubleshooting, and when PandaNpc fits.
Read article →
Connect Claude Code to Your Phone: View Sessions and Approve Tools Anytime on iOS
This article is aimed at developers, sharing best practices for connecting to Claude Code from a mobile phone. Use the PandaNpc iOS app to view sessions in real time, approve tool calls, and respond to questions. With pandapaw and iOS Live Activity, you can achieve efficient remote control and improve coding flexibility.
Read article →