AI Guides › Playbooks

The Three-Task Model Test: Keep, Park or Skip a New AI Model

By Nigel Guy · 6 min read

A new model lands, the launch posts call it the best yet, and you either switch everything on the strength of that or ignore it out of habit. Both feel sensible. Neither tells you whether it handles your work, which is the only question that matters. Leaderboards and demo videos measure someone else's tasks.

The rule: never judge a new model on a fresh, clever prompt. Run it on three tasks you have already done, with a pass mark you wrote down before you pressed enter.

The mechanism: the Test Card

The Test Card is one page, filled in before you open the new model. It takes about twenty minutes, and the time goes on choosing the tasks rather than on running them.

Part What you write Why it matters
Task A: the everyday one A job you do weekly, with a result you already accepted Tests whether it beats your current default on routine work
Task B: the hard one A job where your current tool stumbled or needed reworking Tests the claimed improvement directly
Task C: the trap A job with a missing fact, a false premise or an impossible request Tests whether it admits the gap or bluffs
Pass marks Three to five checks per task, written first Stops you grading on charm
Cost line Time to usable result, plus any usage or price difference A better answer that costs triple may not be better for you

Anthropic's own guidance on testing makes the same points: build tests specific to your use case, include edge cases such as missing or nonexistent input and ambiguous cases, and compare accuracy, response quality and handling of edge cases using your actual prompts and data. OpenAI's evals guide frames it the same way, as specifying how the system should behave before you test it.

Step 1: Pick the three tasks

Choose real work with the sensitive details removed. Keep the source material the same for every model you compare. If you cannot name a task where you know what a good result looks like, you are not ready to judge any model.

Step 2: Write the pass marks

Make them checkable. "Sounds good" is not a mark. "Every figure in the summary appears in the source", "no invented citations", "under 200 words" and "kept my headings" are marks. Three to five per task is plenty.

Step 3: Run the same prompt on both models

Use one prompt for the old and new model, in a fresh conversation each, with the same settings. Where a tool offers a reasoning or effort setting, note what each was on. Anthropic's model guidance says tuning effort is often a better lever than switching models, so a fair comparison fixes it rather than leaving it to defaults. Defaults can differ between models.

Step 4: Score blind if you can

Paste both outputs into a neutral document, label them 1 and 2, shuffle the order, and score against your pass marks before you remember which is which. This is the step people skip, and it removes most of the flattery effect of a new name.

Step 5: Decide with the three-way call

Verdict Condition Action
Keep Passes Task A and B at least as well as your current model, and does not fail Task C worse Move that kind of work over
Park Better on one task, worse on another, or cost or speed is the sticking point Note it on the card, retest at the next release
Skip Fails your marks, or bluffs on the trap Stop. Do not retest for a month

Keep is per task type, not per model. Many people end up with one model for drafting and another for long documents, and that is a perfectly good outcome.

The test prompt

Use this to run each task, so the model's output is easy to score.

You are helping me complete a real piece of work so I can evaluate you fairly.

Context: [CONTEXT_ABOUT_THE_TASK_AND_AUDIENCE]
Source material: [PASTE_SOURCE_MATERIAL]
Task: [WHAT_YOU_WANT_DONE]
What a good result looks like: [YOUR_PASS_MARKS_IN_PLAIN_WORDS]
Format: [LENGTH_AND_STRUCTURE]

Rules:
- Use only the source material and what I have told you here. If a fact you need is missing, say so in a line headed "Missing" instead of filling the gap.
- If any part of the request rests on something that looks wrong or impossible, say so before you start.
- Do not pad. Do not add claims I could not check against the source.

Before you answer, check your draft against each item under "What a good result looks like" and fix any that fail. End with a one-line note listing anything you were unsure of.

Fill in the square-bracket parts once per task. Keep them word-for-word identical across both models.

What it looks like on a real task

A hypothetical example: a freelance bookkeeper wants to know whether a new model suits her client email replies.

Say the new model matches the current one on A, beats it on B, and invents a VAT answer on C. Verdict: keep for tone-sensitive replies, skip for anything involving tax advice, and park the rest. That is a more useful result than "it seems better".

The mistake almost everyone makes

Testing with a prompt you wrote for the old model's quirks, or one so open-ended that any answer looks impressive. The first favours the old model, the second favours whichever writes the most confident prose. A third version: running one task once and treating it as proof. Models vary run to run, so if a result decides something expensive, run it twice.

Also worth knowing: product names and defaults move. The Claude models overview currently lists a lineup that changes with each release, and aliases and pinned IDs behave differently, so note exactly which model name or ID you tested on the card. "The new one" will mean something else in six months.

What to skip

Guardrails

Sources

All 751 AI guides · JulieMango plans from £17/mo