AI Guides › Playbooks
By Nigel Guy · 6 min read
When a new model lands, the first verdict most people reach is about tone: warmer, colder, chattier, drier. It feels like evidence because you can see it in the first three lines. But voice is the cheapest thing to notice and the least useful thing to judge, and Anthropic's own prompting guidance says plainly that "prose style on long-form writing may shift" with a new model, and tells builders to re-check style prompts against the new baseline. A shifted voice is expected. It tells you nothing about whether the model does your job better.
The rule: never judge a model on how it sounds; run the same real job on the old and new model, with a pass mark written down beforehand, and score the output.
At time of writing (4 October 2026), Anthropic's documentation lists Claude Sonnet 5 (released 30 June 2026) as a legacy model, and Claude Sonnet 5.5 (released 28 September 2026) as the current Sonnet. Both are listed at $2 per million input tokens and $10 per million output tokens, which is about £1.50 and about £7.50 at time of writing; check the £ price at checkout or on your cloud bill, because exchange rates move.
We could not verify the claim that Reddit noticed a voice change in Sonnet 5, and we have not tried to. What we can verify is that the official Sonnet 5 prompting guide describes changes that matter more than tone. Everything below comes from that guide.
| Change | What the docs say | Why it affects your test |
|---|---|---|
| Length | Response length is calibrated to task complexity, so shorter on simple lookups and longer on open-ended analysis | A shorter answer is not a worse answer. Score content, not word count |
| Literalism | The model "interprets prompts literally and explicitly" and does not silently generalise an instruction from one item to another | A vague prompt that used to work through goodwill may now return exactly what you typed |
| Effort | Effort defaults to high; Sonnet 5 at medium is described as comparable to Sonnet 4.6 at high |
Compare at matched thinking, not matched labels |
| Agentic behaviour | More likely to use tools and run self-checks by default | Gains show up in multi-step jobs, not one-line chats |
| Thinking | Adaptive thinking is on by default, where Sonnet 4.6 ran without it | Slower or longer replies may be thinking, not padding |
| Tokens | A new tokenizer produces roughly 30% more tokens for the same text | Your cost per job can move even at an unchanged price |
If you use the API, the same guide notes that setting temperature, top_p or top_k to non-default values now returns an error. Sonnet 5.5 has its own migration notes, so read those before moving production code.
Five steps, one page.
1. Pick one real job. Something you did last week and can still check: a report you summarised, a spreadsheet you cleaned, a bug you fixed. A made-up benchmark tells you how the model handles made-up benchmarks.
2. Freeze the inputs. Same files, same prompt, same instructions, saved in one place. Change only the model.
3. Write the pass mark before you run anything. Three to five checks you can answer yes or no. Not "is it good" but "does every figure match the source" and "did it finish without me prompting again".
4. Run both models, twice each. Output varies run to run, so one run each proves little. In the Claude apps, pick the model from the model selector; on the API, set the model ID (claude-sonnet-5 and claude-sonnet-5-5 are the documented IDs). We could not verify which older models your app plan still lists, so check your own selector.
5. Score blind if you can. Paste outputs into a document labelled A and B, with no model names. Score against your checks, then count turns, corrections and time.
Use this prompt to turn your job into checks, or to have a fresh chat score the results blind:
You are a careful reviewer helping me compare two outputs of the same task.
Context: the task was [DESCRIBE THE JOB IN ONE SENTENCE]. The source material is [PASTE OR ATTACH SOURCE].
My pass checks, written before the test:
[LIST 3-5 YES/NO CHECKS]
Below are two outputs, labelled A and B. You do not know which model produced which.
Steps:
1. For each output, answer every check yes or no, quoting the line that justifies the answer.
2. List any factual claim not supported by the source.
3. Note anything the output did that I did not ask for.
4. Give a one-line verdict per output. Do not guess which model wrote it and do not comment on tone unless a check mentions it.
Format: a table with one row per check and columns A and B, then the lists, then the verdicts.
If a check is ambiguous or the source is missing, ask me before scoring rather than guessing. Before you answer, re-read each quote to confirm it exists in the output.
Output A: [PASTE]
Output B: [PASTE]
Fill in the job, source, checks and both outputs.
Say you run a small letting agency and every Monday you turn a messy export of maintenance requests into a tidy list grouped by urgency. Your checks: every request appears once; none are reclassified without a stated reason; the emergency ones come first; no invented addresses; you did not have to send a second message.
Model A produces a friendlier, longer intro and misses two requests. Model B opens flatly, includes all of them, and asks one clarifying question about an ambiguous entry. B scores higher on every check, and it sounds worse to you. That is the point of writing the checks first.
Then re-run B with your prompt tightened for literalism: "Apply the urgency grouping to every request, not just the first ten." If B improves, the gain came from the prompt as much as the model, and that is worth knowing.