Hacker News — AI on Front Page · · 5 min read

Claude Fable 5.1 and Claude Mythos 5.1

Mirrored from Hacker News — AI on Front Page for archival readability. Support the source by reading on the original site.

313 pts · 244 comments on Hacker News

We’re introducing Claude Fable 5.1 and Claude Mythos 5.1. They’re the world’s most advanced models for coding and knowledge work—and their research capabilities offer an early glimpse of how AI models will contribute to scientific progress.

Claude Fable 5.1 and Claude Mythos 5.1 are the same model, but with different levels of safeguards. Fable 5.1 is generally available, while Mythos 5.1 is available only through our trusted access programs; its safeguards are specifically designed to support work in cybersecurity and the life sciences.

Alongside its increased capabilities, Fable 5.1 takes important steps towards addressing the feedback we’ve received from customers on price, data retention, and safeguards.

Price. Fable 5.1 will cost an estimated 25% less than Fable 5 for typical workloads, wherever usage is billed by token. This is because we’re reducing our pricing on cache reads (where the model reads inputs that have already been processed and stored). For highly agentic work, the savings will often be much larger—up to approximately 45%.

Data retention. Our new system of Enterprise Frontier Safeguards (EFS) gives customers complete privacy (the same as a zero data retention policy) while still being state-of-the-art at preventing adversarial use. EFS works by storing data in cloud infrastructure controlled entirely by the customer, not Anthropic. It will be made available to enterprise customers in phases, beginning later this fall. Until EFS is available, eligible customers will be able to use Fable 5.1 with zero data retention.

Safeguards. We’ve improved our safeguards to reduce false positives (where the system flags benign content). In cybersecurity, our newest safeguards block 60% fewer false positives than before. In part, this is because Fable 5.1 can now be used to discover software vulnerabilities—though not develop exploits for them. In biology, we’ve established an access program, developed in partnership with the US government, to enable access to Claude Mythos 5.1’s advanced biology capabilities. We expect to open enrollment for scientists soon.

A new performance frontier

Claude Fable 5.1 sets a new standard on coding, knowledge work, and long-running problem-solving tasks. The charts below show that Fable 5.1 is capable of much higher performance than its predecessor, Fable 5. And when set to Low or Medium effort, Fable 5.1 achieves similar or better results than Fable 5 at a much lower cost. (Note that Fable 5.1 defaults to High effort in Claude Code, and to Medium in Claude Cowork and on Claude.ai.)

Agentic scientific researchAgentic terminal codingMultidisciplinary reasoningAgentic coding
Terminal-Bench-Science 0.1Accuracy vs Cost
  • Fable 5.1
  • Fable 5

Terminal-Bench-Science 0.1 scores by cost (log scale), at each effort level.

Terminal-Bench 4.0Accuracy vs Cost
  • Mythos 5.1
  • Fable 5.1
  • Mythos 5

Terminal-Bench 4.0 scores by cost (log scale), at each effort level. Claude Fable 5.1 and Claude Mythos 5.1 are the same underlying model; the gap between them reflects the tasks on which our earlier, less precise cyber safeguards intervened. With the improvements we’re making to these safeguards today, we expect the difference between the models to be much smaller.

Humanity's Last ExamAccuracy vs Cost
  • Fable 5.1 (with tools)
  • Fable 5.1 (no tools)
  • Fable 5 (with tools)
  • Fable 5 (no tools)

Humanity’s Last Exam scored by cost (log scale), at each effort level.

CursorBench 3.2.0Accuracy vs Cost
  • Fable 5.1
  • Fable 5

CursorBench 3.2.0 by cost (log scale), at each effort level.

Fable 5.1 avoids shortcuts that result in poorer-quality work, and it’s smart enough to fix the root causes of software issues. For example, in testing by the investment firm Millennium, Fable 5.1 found the cause of a rare crash on their internal systems that none of their engineers (or any other model) had been able to explain after several years of trying.

Here, you can see how Fable 5.1 compares across various benchmarks:

Fable 5.1Fable 5Opus 5GPT-5.6 Sol
Agentic scientific researchTerminal-Bench-Science 0.1
Agentic scientific researchTerminal-Bench-Science 0.152.6%24.7%29.0%22.4%
Agentic codingTerminal-Bench 4.0
Agentic codingTerminal-Bench 4.055.8%60.9% (Mythos 5.1)42.0%52.3%37.3%
Knowledge workGDPval-AA v2
Knowledge workGDPval-AA v21853172318241711
Computer useOSWorld 2.0
Computer useOSWorld 2.077.9%partial72.9%partial75.4%partialpartial
41.7%strict36.1%strict39.6%strictstrict
Multidisciplinary reasoningHumanity's Last Exam
Multidisciplinary reasoningHumanity's Last Exam60.9%no tools57.8%no tools56.6%no toolsno tools
65.0%with tools63.8%with tools63.6%with toolswith tools
Business workflowsAutomationBench
Business workflowsAutomationBench31.4%17.1%26.9%19.6%
Agentic codingCursorBench 3.2.0
Agentic codingCursorBench 3.2.073.4%70.5%70.0%67.2%
Fable 5.1 was evaluated with its production safeguards enabled. On tasks where these safeguards intervened, Fable 5.1 and Fable 5 scored a zero on OSWorld 2.0, and Fable 5 scored a zero on AutomationBench. In all other interventions from our safeguards, cybersecurity tasks were completed by Claude Opus 4.8, and biology tasks were completed by Claude Opus 5. This likely reduces the performance of Fable 5.1 and Fable 5 on these benchmarks. Terminal-Bench-Science 0.1: The standard error is ±3.5–4.5 pts per model. The public leaderboard (3 trials/task, Claude Code harness) reports Claude Opus 5 at 30.0% and Claude Fable 5 at 21.4%; our setup reproduces them at 29.0% and 24.7%, respectively, both within noise. OSWorld 2.0: Scores are on the benchmark authors’ August 2026 task release; Fable 5 and Opus 5 were re-run under the same conditions. Because the task files differ from earlier releases, these numbers aren't directly comparable to previously published OSWorld 2.0 results, which is why no competitor score is shown.

Our early-access partners noticed these performance upgrades, and also picked up on more qualitative improvements in the model’s outputs. Here’s what they told us:

Quote

“In internal benchmarks, Claude Fable 5.1 solves more of our coding problems than Fable 5 or Opus 5, and achieves state of the art on trading intuition. While prior models became hard to follow the longer they worked, Fable 5.1 remains readable over long, multi-step tasks.”

CompanyJane Street Capital
AuthorCraig Falls, Head of Quantitative Research

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hacker News — AI on Front Page