AI Guides › Playbooks

The Same-Job Scorecard: Judging a New Claude Model by the Work, Not the Voice

By Nigel Guy · 6 min read

When a new model lands, the first verdict most people reach is about tone: warmer, colder, chattier, drier. It feels like evidence because you can see it in the first three lines. But voice is the cheapest thing to notice and the least useful thing to judge, and Anthropic's own prompting guidance says plainly that "prose style on long-form writing may shift" with a new model, and tells builders to re-check style prompts against the new baseline. A shifted voice is expected. It tells you nothing about whether the model does your job better.

The rule: never judge a model on how it sounds; run the same real job on the old and new model, with a pass mark written down beforehand, and score the output.

What changed, and what we could not confirm

At time of writing (4 October 2026), Anthropic's documentation lists Claude Sonnet 5 (released 30 June 2026) as a legacy model, and Claude Sonnet 5.5 (released 28 September 2026) as the current Sonnet. Both are listed at $2 per million input tokens and $10 per million output tokens, which is about £1.50 and about £7.50 at time of writing; check the £ price at checkout or on your cloud bill, because exchange rates move.

We could not verify the claim that Reddit noticed a voice change in Sonnet 5, and we have not tried to. What we can verify is that the official Sonnet 5 prompting guide describes changes that matter more than tone. Everything below comes from that guide.

Where the upgrade actually is

Change What the docs say Why it affects your test
Length Response length is calibrated to task complexity, so shorter on simple lookups and longer on open-ended analysis A shorter answer is not a worse answer. Score content, not word count
Literalism The model "interprets prompts literally and explicitly" and does not silently generalise an instruction from one item to another A vague prompt that used to work through goodwill may now return exactly what you typed
Effort Effort defaults to high; Sonnet 5 at medium is described as comparable to Sonnet 4.6 at high Compare at matched thinking, not matched labels
Agentic behaviour More likely to use tools and run self-checks by default Gains show up in multi-step jobs, not one-line chats
Thinking Adaptive thinking is on by default, where Sonnet 4.6 ran without it Slower or longer replies may be thinking, not padding
Tokens A new tokenizer produces roughly 30% more tokens for the same text Your cost per job can move even at an unchanged price

If you use the API, the same guide notes that setting temperature, top_p or top_k to non-default values now returns an error. Sonnet 5.5 has its own migration notes, so read those before moving production code.

The mechanism: the Same-Job Scorecard

Five steps, one page.

1. Pick one real job. Something you did last week and can still check: a report you summarised, a spreadsheet you cleaned, a bug you fixed. A made-up benchmark tells you how the model handles made-up benchmarks.

2. Freeze the inputs. Same files, same prompt, same instructions, saved in one place. Change only the model.

3. Write the pass mark before you run anything. Three to five checks you can answer yes or no. Not "is it good" but "does every figure match the source" and "did it finish without me prompting again".

4. Run both models, twice each. Output varies run to run, so one run each proves little. In the Claude apps, pick the model from the model selector; on the API, set the model ID (claude-sonnet-5 and claude-sonnet-5-5 are the documented IDs). We could not verify which older models your app plan still lists, so check your own selector.

5. Score blind if you can. Paste outputs into a document labelled A and B, with no model names. Score against your checks, then count turns, corrections and time.

Use this prompt to turn your job into checks, or to have a fresh chat score the results blind:

You are a careful reviewer helping me compare two outputs of the same task.

Context: the task was [DESCRIBE THE JOB IN ONE SENTENCE]. The source material is [PASTE OR ATTACH SOURCE].

My pass checks, written before the test:
[LIST 3-5 YES/NO CHECKS]

Below are two outputs, labelled A and B. You do not know which model produced which.

Steps:
1. For each output, answer every check yes or no, quoting the line that justifies the answer.
2. List any factual claim not supported by the source.
3. Note anything the output did that I did not ask for.
4. Give a one-line verdict per output. Do not guess which model wrote it and do not comment on tone unless a check mentions it.

Format: a table with one row per check and columns A and B, then the lists, then the verdicts.

If a check is ambiguous or the source is missing, ask me before scoring rather than guessing. Before you answer, re-read each quote to confirm it exists in the output.

Output A: [PASTE]
Output B: [PASTE]

Fill in the job, source, checks and both outputs.

A worked example (hypothetical)

Say you run a small letting agency and every Monday you turn a messy export of maintenance requests into a tidy list grouped by urgency. Your checks: every request appears once; none are reclassified without a stated reason; the emergency ones come first; no invented addresses; you did not have to send a second message.

Model A produces a friendlier, longer intro and misses two requests. Model B opens flatly, includes all of them, and asks one clarifying question about an ambiguous entry. B scores higher on every check, and it sounds worse to you. That is the point of writing the checks first.

Then re-run B with your prompt tightened for literalism: "Apply the urgency grouping to every request, not just the first ten." If B improves, the gain came from the prompt as much as the model, and that is worth knowing.

What to skip

Guardrails

Sources

All 751 AI guides · JulieMango plans from £17/mo