Fable 5.1 vs. Fable 5: Results on a real-world budget, not the spec sheet
The New Stack

Fable 5.1 vs. Fable 5: Results on a real-world budget, not the spec sheet

Fable 5.1 vs. Fable 5: Results on a real-world budget, not the spec sheet When Anthropic launched Claude Fable 5.1 this month, it centered the announcement around one benchmark result: its Terminal-Bench-Science score. In this benchmark, a model gets a terminal and a real scientific research problem to solve independently. Fable 5.1 scores 52.6%, and Fable 5 scores 24.7%. By Anthropic’s scoring, the new model more than doubles the old one. Anthropic’s published score was produced under conditions most users don’t have access to. The benchmark allows each model up to eight hours per task, and Anthropic has not said what harness or budget it used to get its numbers. When the benchmark’s leaderboard tested Fable 5, it ran the model through Claude Code at maximum effort and spent $14,180 across 210 attempts, about $67 each. Most people don’t use Fable in a lab, so I wanted to know what an average user would get on these tasks. I recently tested Fable 5 and Fable 5.1 on everyday work and found them far closer than the benchmark suggests. That made me want to run the benchmark’s own tasks the way a home user would and see where the models actually differ. The benchmark’s 70 tasks are public, so I pulled five of them, one from each science field, and ran both models myself. The tests Terminal-Bench-Science has five categories, each with multiple tests. I chose one test per category that could run in a Python environment. Here’s what I picked: - Symbolic regression (mathematics) - A dataset with 100 variables and a hidden formula behind a yes-or-no label. The model must find a predictor that works on data it has never seen. - Lorenz-96 assimilation (Earth sciences) - Reconstruct a chaotic atmospheric model from a few uncalibrated sensors with unknown clock offsets. Grading is all-or-nothing on five criteria. - Reactor safety control (engineering) - Write a controller for a chemical reactor that finishes every batch as fast as possible without ever exceeding the temperature limit, across public and hidden fault scenarios. - Foraging cognitive model (life sciences) - Predict, trial by trial, which lever each of 20 mice will press, graded on sessions the model never saw. - Nanoindentation (physical sciences) - Extract material properties from raw indentation curves that include drift, adhesion, defects, and an unknown tip shape. Each run got a plain terminal, and I set a $12 limit and 60 turns for each test. The full set of ten runs took about 12 hours. Symbolic regression This was the only test where a model passed the benchmark’s hidden test. Fable 5.1 worked for 27 turns, found the hidden structure, wrote a predictor, and stopped on its own after 11.8 minutes, 27,088 output tokens, and cost $1.96 to pass this one test. Fable 5 used all 60 turns over 53.5 minutes, generated 39,461 output tokens, cost $4.20, and failed. I ran Fable 5 a second time to rule out bad luck. It used all 60 turns again, took 60 minutes, generated 60,608 output tokens, cost $6.38, and failed again. Lorenz-96 assimilation This was the most expensive pair of runs. Fable 5 hit the $12 cost limit at 45 turns after 97.7 minutes and 92,091 output tokens, ending at $12.63. Fable 5.1 used all 60 turns over 126 minutes, generated 89,789 output tokens, and cost $10.70. On the public leaderboard, Earth sciences is also the field where Fable 5 scores close to zero, and both models failed it here. Reactor safety control Neither wrote a controller that passed the grader’s scenarios. This run produced the most output tokens, 157,710 generated by Fable 5.1. It hit the 60-turn limit after 40.9 minutes and cost $11.53. Fable 5 hit the cost limit at 49 turns after 63.8 minutes, 121,978 output tokens, and $12.04. Foraging cognitive model This was the longest run of the testing series. Fable 5.1 was the only model that declared itself finished. It built a model, tested it against its own scoring loop, and declared it done at 43 turns after 53.5 minutes. It created 65,518 output tokens and cost $5.65. But the official grader rejected it. Fable 5 never declared anything. It hit the $12 limit at 60 turns, after 139.3 minutes and 62,587 output tokens, ending at $12.13. Nanoindentation Both failed. Both spent most of the run reading raw curves and writing code to segment them. Neither produced a results file the grader accepted. Fable 5.1 ran out of turns at 29.6 minutes, 115,687 output tokens, and $10.91. Fable 5 ran out of money at 48 turns after 34.1 minutes and 114,239 output tokens, ending at $12.59. Results Here are the results by the numbers. | Metric | Fable 5 score | Fable 5.1 score | | Tasks solved | 0 of 5 | 1 of 5 | | Output tokens | 430,356 | 455,792 | | Total cost | $53.59 | $40.75 | | Total time | 388 min | 262 min | | Runs ended by cost limit | 4 | 0 | The benchmark scores models on all 70 tasks with three trials each, and Anthropic’s 24.7% and 52.6% come from that full suite. The independent leaderboard puts Fable 5 at 21.4%, close to Anthropic’s figure. Fable 5.1 is not on the independent leaderboard yet, so its 52.6% is Anthropic’s number alone. My results, 0% and 20%, are below both. Five tasks are a small sample. Getting these results by chance is plausible even if the published scores are exactly right, so this run neither confirms nor contradicts the doubling claim. The direction matched, since the new model did better. The one task Fable 5.1 solved was in mathematics, which is also the field where the leaderboard shows Fable 5 performing best. What I think I don’t think a regular user will see much difference between Fable 5 and Fable 5.1. I only ran a small sample of tests, so I can’t prove or reject Anthropic’s benchmark results. But what I saw suggests that gap won’t reach the average user. The one difference that did show up was on the bill. Fable 5.1 failed faster and cheaper, and it never hit my cost limit, whereas Fable 5 hit it four times. Suppose your work looks more like the benchmark tasks; the harness and the budget matter as much as the model. With a purpose-built harness, hours per task, and a much bigger budget, you may get closer to Anthropic’s numbers.

Read on The New Stack ↗ ← Back to News

Comments

No comments yet. Start the discussion.