AI Guides › Workbench
By Nigel Guy · 7 min read
Most people meet a new model the way they meet a new phone: they glance at the launch headlines, ask it something they already know the answer to, see a decent reply, and carry on. Claude Fable 5.1 was released on 1 September 2026 and plenty of people did exactly that. A familiar question tells you almost nothing, because the old model already handled it. The only useful test is one that sits at the edge of what you actually do.
The rule: test a new model on your own real work, one narrow trial at a time, and decide each trial's pass mark before you run it.
This kit has four trials: cost, honesty, audit and unattended running. Each has a prompt you can paste today.
Anthropic's own pages confirm that Claude Fable 5.1 exists, was released on 1 September 2026, and is its most capable generally available model. It follows Fable 5 (June 2026). Mythos 5.1 is the same model with different safeguards, for vetted cyber and life-science users only. Fable 5.1 sits at the top of the range, above Opus 5.5 and Sonnet 5.5, and Anthropic says it suits long-running, asynchronous work.
We did not independently test it. The benchmark figures and customer quotes in Anthropic's announcement are Anthropic's and its partners' claims, and we use none of them here. The kit exists so you can generate your own evidence instead.
| Trial | What it tests | Cost at time of writing | Best for | Catch |
|---|---|---|---|---|
| 1. Effort Pair | Whether lower effort is good enough for your task | Counts against usage; see below | Anyone paying per token or watching limits | One task is one data point |
| 2. Blunt Read | Whether it will disagree with you | Same | Stuck decisions and plans | You still decide; it has no stake |
| 3. Inheritance Audit | Whether it finds quiet damage | Same | Code, documents, processes you built | Can overstate; verify each finding |
| 4. Small Unattended Agent | Whether it can finish alone | Highest usage of the four | Narrow, repeatable jobs | Needs tight permissions |
Fable 5.1 is on paid Claude plans only. According to Anthropic's help centre, on Max plans (and premium seats on Team and seat-based Enterprise) it is part of the plan, and you can use up to 50% of your weekly limits on Fable models. On Pro and standard Team seats it is not included in plan limits and runs on pay-as-you-go usage credits. Check your own plan page, because this has changed before: Fable 5 had a promotion that ended in July.
On the API, Anthropic lists $10 per million input tokens and $50 per million output tokens (about £7.50 and £37 at time of writing; check the £ price on your own bill, as exchange rates move). Cache reads are $0.25 per million, which is 75% less than on Fable 5. In claude.ai, pick it from the model picker. In Claude Code you need version 2.1.255 or later. Defaults differ: Anthropic says it defaults to High effort in Claude Code and Medium in Cowork and on claude.ai.
In claude.ai, click the model name next to the send button, then Effort, and choose a level. Levels are Low, Medium, High, Extra high and Max. Higher effort gives more thorough answers but is slower and eats usage faster. Do not rely on the model to switch levels itself; run the same task twice by hand, changing only the setting.
Pass mark, written first: "If the Low or Medium answer would have been sent without edits, it passes."
ROLE: You are a careful reviewer helping me decide which effort setting is worth paying for.
CONTEXT: Below is a real task from my work. I am running it twice in separate chats, once at a low effort setting and once at a high one. I will paste both outputs back to you.
TASK: [PASTE ONE REAL TASK, WITH ALL THE INPUTS IT NEEDS]
WHAT GOOD LOOKS LIKE: [YOUR PASS MARK, e.g. "usable without edits"]
Step 1: If any input is missing or ambiguous, ask me before answering.
Step 2: Complete the task.
Step 3: End with a three-line note: what you assumed, what you are least sure of, and what a harder look might have changed.
Do not pad. Do not claim certainty you lack.
Then paste both outputs into a fresh chat and ask for a plain comparison against your pass mark. Judge the difference yourself first.
Pick the problem you have been circling. Include what you have already tried; without it you get the advice you have already rejected.
ROLE: You are a direct, well-informed adviser with no interest in flattering me.
SITUATION: [DESCRIBE THE PROBLEM IN PLAIN TERMS]
ALREADY TRIED: [WHAT YOU DID AND WHAT HAPPENED]
CONSTRAINTS: [TIME, MONEY, PEOPLE, THINGS YOU WON'T DO]
Respond in this order:
1. What you think is really going on, in two or three sentences.
2. What I may be avoiding or underweighting.
3. The single next step you would take, and why that one first.
4. One thing you would need to know from me to be more confident.
Ask me questions first if the situation is unclear. Do not soothe, and do not invent facts about my circumstances. Before answering, check that you have not simply agreed with me.
Pass mark: it tells you at least one thing you did not already know, and it disagrees somewhere. Pure agreement is a fail.
ROLE: You are a maintainer about to take over what I built below, and you will be on call for it.
THING TO AUDIT: [PASTE CODE, DOCUMENT OR PROCESS, PLUS WHAT IT IS MEANT TO DO]
HOW IT IS USED: [WHO RELIES ON IT, HOW OFTEN]
Find real defects and fragilities only: things that could fail, mislead or corrupt without anyone noticing. Skip style opinions.
Output a table: issue, where it is, how it fails, how visible the failure is (loud or silent), severity (high, medium, low). Order by silent damage first.
Then name the one fix you would do first and why.
For each issue, quote the exact line or sentence you are relying on. If you cannot point to it, mark the issue "unverified". If you need more context, ask before assessing.
Pass mark: at least one finding you confirm is real. Check every "high" yourself before acting.
Anthropic positions Fable 5.1 for jobs that run for hours, but start with something that runs for minutes. Use Claude Code or Cowork, give it access to a copy of the data and nothing else, and stay near the first run.
ROLE: You are an assistant running one small, repeatable job without me watching.
JOB: [ONE NARROW TASK]
INPUTS YOU MAY USE: [FOLDERS, FILES OR SOURCES]
OFF LIMITS: [ANYTHING YOU MUST NOT CHANGE, DELETE, SEND OR SPEND]
Before you start, write back: your definition of "done" in checkable terms, what you will do if you get stuck (stop and write down the blocker; never improvise outside the inputs), and what record you will leave.
Then do the job. When finished, leave a short log: what you did, what you skipped and why, anything you are unsure about, and how I can verify the result in under five minutes.
Ask me for anything missing before you begin. If a step needs something off limits, stop and say so.
Pass mark: you can verify the result quickly, and it stopped rather than guessing when blocked. Plant one deliberate obstacle to see which.