AI Guides › Playbooks
By Nigel Guy · 6 min read
A new model lands, the launch posts call it the best yet, and you either switch everything on the strength of that or ignore it out of habit. Both feel sensible. Neither tells you whether it handles your work, which is the only question that matters. Leaderboards and demo videos measure someone else's tasks.
The rule: never judge a new model on a fresh, clever prompt. Run it on three tasks you have already done, with a pass mark you wrote down before you pressed enter.
The Test Card is one page, filled in before you open the new model. It takes about twenty minutes, and the time goes on choosing the tasks rather than on running them.
| Part | What you write | Why it matters |
|---|---|---|
| Task A: the everyday one | A job you do weekly, with a result you already accepted | Tests whether it beats your current default on routine work |
| Task B: the hard one | A job where your current tool stumbled or needed reworking | Tests the claimed improvement directly |
| Task C: the trap | A job with a missing fact, a false premise or an impossible request | Tests whether it admits the gap or bluffs |
| Pass marks | Three to five checks per task, written first | Stops you grading on charm |
| Cost line | Time to usable result, plus any usage or price difference | A better answer that costs triple may not be better for you |
Anthropic's own guidance on testing makes the same points: build tests specific to your use case, include edge cases such as missing or nonexistent input and ambiguous cases, and compare accuracy, response quality and handling of edge cases using your actual prompts and data. OpenAI's evals guide frames it the same way, as specifying how the system should behave before you test it.
Choose real work with the sensitive details removed. Keep the source material the same for every model you compare. If you cannot name a task where you know what a good result looks like, you are not ready to judge any model.
Make them checkable. "Sounds good" is not a mark. "Every figure in the summary appears in the source", "no invented citations", "under 200 words" and "kept my headings" are marks. Three to five per task is plenty.
Use one prompt for the old and new model, in a fresh conversation each, with the same settings. Where a tool offers a reasoning or effort setting, note what each was on. Anthropic's model guidance says tuning effort is often a better lever than switching models, so a fair comparison fixes it rather than leaving it to defaults. Defaults can differ between models.
Paste both outputs into a neutral document, label them 1 and 2, shuffle the order, and score against your pass marks before you remember which is which. This is the step people skip, and it removes most of the flattery effect of a new name.
| Verdict | Condition | Action |
|---|---|---|
| Keep | Passes Task A and B at least as well as your current model, and does not fail Task C worse | Move that kind of work over |
| Park | Better on one task, worse on another, or cost or speed is the sticking point | Note it on the card, retest at the next release |
| Skip | Fails your marks, or bluffs on the trap | Stop. Do not retest for a month |
Keep is per task type, not per model. Many people end up with one model for drafting and another for long documents, and that is a perfectly good outcome.
Use this to run each task, so the model's output is easy to score.
You are helping me complete a real piece of work so I can evaluate you fairly.
Context: [CONTEXT_ABOUT_THE_TASK_AND_AUDIENCE]
Source material: [PASTE_SOURCE_MATERIAL]
Task: [WHAT_YOU_WANT_DONE]
What a good result looks like: [YOUR_PASS_MARKS_IN_PLAIN_WORDS]
Format: [LENGTH_AND_STRUCTURE]
Rules:
- Use only the source material and what I have told you here. If a fact you need is missing, say so in a line headed "Missing" instead of filling the gap.
- If any part of the request rests on something that looks wrong or impossible, say so before you start.
- Do not pad. Do not add claims I could not check against the source.
Before you answer, check your draft against each item under "What a good result looks like" and fix any that fail. End with a one-line note listing anything you were unsure of.
Fill in the square-bracket parts once per task. Keep them word-for-word identical across both models.
A hypothetical example: a freelance bookkeeper wants to know whether a new model suits her client email replies.
Say the new model matches the current one on A, beats it on B, and invents a VAT answer on C. Verdict: keep for tone-sensitive replies, skip for anything involving tax advice, and park the rest. That is a more useful result than "it seems better".
Testing with a prompt you wrote for the old model's quirks, or one so open-ended that any answer looks impressive. The first favours the old model, the second favours whichever writes the most confident prose. A third version: running one task once and treating it as proof. Models vary run to run, so if a result decides something expensive, run it twice.
Also worth knowing: product names and defaults move. The Claude models overview currently lists a lineup that changes with each release, and aliases and pinned IDs behave differently, so note exactly which model name or ID you tested on the card. "The new one" will mean something else in six months.