AI Guides › Workbench

The Seven-Day Model Log: Testing a New Claude Model Before You Trust It

By Nigel Guy · 7 min read

Most "week with the new model" posts are a day-one impression stretched over seven days: a few impressive demos, one disappointing answer, a verdict that matches whatever the writer expected. It feels like testing because you used the thing a lot. It isn't, because nothing was written down before the model got a chance to flatter you.

The rule: write the tasks and the pass mark before day one, log every real use as it happens, and give a verdict only from the log.

One honest note first. This guide gives you the method and a clearly hypothetical sample log. It does not contain a week of our own measurements, and it does not claim to. Claude Sonnet 5.5 was released on 28 September 2026, so at the time of writing it is about a week old, and any confident "what held up" claim from anyone is early.

What is actually current

Check this before you test, because "Sonnet 5" is already the previous model in this line.

Anyone writing "Sonnet 5" reviews in October is probably describing a model that has been superseded. Name the exact model in your log.

Before day one

You need Detail
A model you can name exactly In the app, check the model picker. In the API, the ID is claude-sonnet-5-5.
Access Free plan works for Sonnet 5.5. Pro is $20 a month, about £15 at time of writing (check the £ price at checkout). Max starts from $100, about £75.
API pricing, if you use it $2 per million input tokens and $10 per million output tokens, about £1.50 and about £7.50 at time of writing. Batch requests are half price.
Ten real tasks Work you would do anyway this week, not puzzles.
A baseline The model you use today, or your own result, for the same tasks.
A spreadsheet or notes file One row per use.

The kit

Tool What it does Cost Best for Catch
Test Card Fixes tasks, pass marks and baseline before you start Free Stopping yourself moving the goalposts Takes thirty minutes you will be tempted to skip
Daily log One row per real use, written within the hour Free Catching patterns a day-seven memory forgets Only works if you log failures too
Side-by-side run Same prompt on the old and new model Free to a small usage cost Telling "better" from "different" Two runs per prompt; one run proves little
Cost meter Tracks tokens or plan usage per task Free in the Console; plan usage limits apply on the apps Spotting cost drift Effort settings change thinking and cost
Verdict sheet Turns the log into one of three decisions Free Ending the test You must accept a boring answer

Step 1: Build the Test Card (day zero)

List ten tasks. For each, write what a pass looks like, what a fail looks like, and your baseline result. Fix the mix: roughly a third routine, a third hard-for-you, a third where you can check facts yourself. Use this prompt to draft the card.

You are helping me design a fair one-week test of an AI model for my own work.

Context: I am a [YOUR_ROLE] and I will test [MODEL_NAME_AND_VERSION] against [BASELINE_MODEL_OR_METHOD].
My typical work this week: [LIST_OF_REAL_TASKS_OR_PROJECTS].

Goal: a Test Card of exactly ten tasks drawn from my real work, with a mix of routine, difficult, and fact-checkable tasks.

For each task give: a one-line description; the input I must supply; a pass condition I can check in under five minutes; a fail condition; and which baseline I compare against.

Rules:
- Do not invent details about my work. If anything above is missing or vague, ask me questions first and wait.
- Pass conditions must be observable (a number, a checklist, a diff), never "feels good".
- Include two tasks designed to expose confident mistakes, where I already know the right answer.
- Output as a table, then three lines on what this card cannot tell me.

Before answering, check that every pass condition could be judged by someone who has not seen the output.

Fill in: your role, the model and baseline names, and your real task list.

Step 2: Log every use (days one to seven)

One row per use: date, task ID, effort setting if you can see it, pass or fail, minutes saved or lost against the baseline, and one sentence on what went wrong. Log failures first. Three columns matter most: did it pass, did you have to fix it, and would you have caught the error without the check you built in.

Step 3: Run side by side on the hard tasks

For the three tasks you care about most, run the same prompt on the previous model and the new one, twice each. Anthropic's docs say effort levels differ between Sonnet 5 and 5.5, so if you use the API, record the effort level rather than assuming the default matches.

Step 4: Write the verdict from the log

You are an analyst reviewing my week-long model test. Use only the log below; do not add outside knowledge about the model's quality.

Test Card: [PASTE_TEST_CARD]
Daily log: [PASTE_LOG_ROWS]

Produce:
1. Pass rate by task type (routine, difficult, fact-checkable).
2. The three most costly failures, quoting my own log notes.
3. Any pattern across failures, stated only if it appears at least twice.
4. One verdict from: Switch, Split, Stay. Give the single strongest reason from the log.
5. Two things the log cannot tell us.

If the log has fewer than [MINIMUM_ROWS] rows or no failures recorded, say the evidence is too thin and ask what is missing instead of giving a verdict.
Before answering, confirm every claim points to a log row.

A hypothetical week

This is invented to show the format, not a result. Suppose a freelance copywriter tests ten tasks.

Day Task Result Note
1 Rewrite a product page Pass Slightly faster than baseline
2 Summarise a 40-page PDF Pass One figure misquoted, caught on check
3 Client email in my voice Fail Too formal, two rounds to fix
5 Spreadsheet formula debug Pass Correct first time
6 Fact-check trap Fail Confident wrong date

Two fails in five rows is a pattern about verification, not a verdict on the model. That is why you need ten tasks and a pass mark.

The trap

Novelty. The first two days feel excellent because everything is new and your prompts are fresh. Late in the week you reuse the same prompts and the gaps appear. Weight days five to seven as heavily as days one to three, and never score a task you did not write down in advance.

The three verdicts

Verdict Use when
Switch It passes at least as often as the baseline, and your failures are cheaper to catch
Split It wins on some task types and loses on others; route by type
Stay No clear gain, or the failures are ones you would not catch

"Not enough data, extend the test" is allowed once. Twice is avoidance.

What to skip

Guardrails

Sources

All 751 AI guides · JulieMango plans from £17/mo