AI Guides › Playbooks

Three Tests on Your Own Work: The Model Comparison Card

By Nigel Guy · 7 min read

Most people pick an AI model the way they pick a restaurant: a leaderboard, a friend's recommendation, or one impressive answer to one clever question. That feels like research while it is happening. It is not, because a benchmark measures somebody else's tasks, and one good answer tells you about one afternoon.

The rule: never choose a model on public rankings or a single answer. Run the same three tests on your own work, score them against criteria you wrote before you saw any output, and let the card decide.

The mechanism: the Model Comparison Card

The card is one page. It has a row per test and a column per model. You fill in the criteria first and the scores second. Nothing else goes on it.

The shape follows how the major labs describe their own evaluation advice. Anthropic's testing guidance says to define success criteria that are specific and measurable, to mirror the real task, to include edge cases, and to expect more than one dimension (accuracy, tone, consistency, latency and price all count). OpenAI's evals guide describes the same loop: describe the task, run it with representative inputs, analyse, iterate. You are doing a small, manual version of that.

Step 0: write the criteria before you run anything

Pick three to five things that make an output good for you. Examples: facts match my source; tone sounds like me; no invented details; format I can paste straight in; under a length limit. Give each a pass/fail or a 1 to 3 score. If you write the criteria after reading the outputs, you will unconsciously write the criteria that the nicest-sounding output already meets.

Test 1: The Replay

Take five to ten real tasks you have already done, where you know what a good result looks like. A reply to a difficult client, a summary of a document you read closely, a spreadsheet formula you had to fix. Remove anything confidential or check the tool's data settings first (see Guardrails).

Run each task on each model with the same wording, in a fresh conversation. Score against your criteria. Because you already know the answer, you are judging against evidence rather than against how confident the output sounds.

Test 2: The Blind Side-by-Side

Judging by brand name is a bias. Copy each model's output into a document, label them A, B and C in random order, and strip anything that gives the source away (greetings, signature phrases, formatting tics you could recognise). Score or rank them against the criteria. Only then reveal which was which.

If you have a colleague, have them do the ranking and keep your own scores to compare. If you use another model as the grader, use one that is not in the contest, because Anthropic's guidance recommends a separate evaluator model for model-graded scoring. Treat a model grader as a first pass; your own read of the winner still counts.

Test 3: The Stress Run

The Replay uses tidy inputs. Real work is not tidy. Build five awkward cases: a document with a contradiction in it, an instruction that is ambiguous, an input that is far longer than usual, a request where the right answer is "I can't tell from this", and a request containing a detail that is nearly right but wrong. Anthropic's list of edge cases covers similar ground: irrelevant or missing input, overly long inputs, ambiguity, and conflicting information.

Score two things: did it handle the trap, and did it say so when it was unsure. A model that confidently fills the gap with invention fails, however fluent it sounds. Then repeat two of your Replay tasks a second time. If the same model gives you a good result once and a poor one the next time, that inconsistency belongs on the card.

A worked example (hypothetical)

Imagine Priya, who runs a small bookkeeping practice and wants help drafting client emails and summarising HMRC letters. She tests two assistants.

Criteria, written first: every figure matches the source letter; plain English a client would follow; no advice beyond the letter; under 150 words for emails.

Test Assistant A Assistant B
Replay: 6 past letters, figures correct 5 of 6 6 of 6
Replay: plain English (1 to 3 average) 2.8 2.0
Blind side-by-side: Priya's ranking 1st on 4 of 6 1st on 2 of 6
Stress: spotted the contradiction in the letter Yes, flagged it No, picked one figure
Stress: said "can't tell" when asked about a missing year Yes Invented a figure
Repeat run consistency Same result One of two repeats wrong

Assistant B's prose might win a leaderboard. A is the better tool for Priya, because her work punishes invention. The numbers here are invented to show the layout; they are not results from any real product.

A prompt to build your test set

Use this to turn a description of your work into awkward cases quickly. Fill in the square-bracket parts.

You are helping me design a fair test for comparing AI models on my real work.

My job and context: [YOUR_ROLE_AND_CONTEXT]
The task I want to test: [TASK_DESCRIPTION]
What a good result looks like: [SUCCESS_CRITERIA]
Typical input I deal with: [DESCRIBE_OR_PASTE_ANONYMISED_EXAMPLE]

If any of those four are missing or vague, ask me questions first and do not guess.

Then produce:
1. Five realistic test inputs of normal difficulty, each with a one-line note on what a correct answer must contain.
2. Five stress inputs, one for each of: a contradiction inside the input, an ambiguous instruction, an unusually long input, a case where the honest answer is "this can't be determined", and a detail that is almost right but wrong. Mark each with the trap it contains and the correct handling.
3. A scoring sheet as a table, with my criteria as columns and a 1 to 3 scale defined in plain words.

Constraints: use only invented or generic details, never real names or figures. Keep every test input short enough to paste into a chat. Before answering, check that each stress input really contains the trap you labelled and that no two tests overlap.

A prompt for the blind grader

Only use this with a model that was not one of the contestants, and paste outputs under neutral labels.

You are a careful, impartial reviewer. Below is a task, my success criteria, and [NUMBER] anonymous responses labelled [LABELS].

Task: [TASK_TEXT]
Source material the responses should rely on: [SOURCE_TEXT_OR_NOTES]
My criteria: [CRITERIA_LIST]

For each response, score each criterion 1 to 3 and give one sentence of evidence quoted or closely paraphrased from the response. Flag any claim that is not supported by the source material. Do not reward length, confidence or polish for their own sake. Then rank the responses and state how close the call is. If the source material is missing or insufficient to judge a criterion, say so instead of guessing.

Before answering, check that your scores match your evidence sentences and that your ranking follows from the totals.
Output as a table, then the ranking, then a two-sentence summary.

What it costs

Most chat assistants have a free tier, so you can often run Test 1 for nothing. Claude's page at time of writing lists Free at $0, Pro at $20 a month (about £16 at time of writing; check the £ price at checkout) and Max from $100 a month. A paid month for one or two contenders is usually enough for a proper comparison. Check each vendor's current pricing page, as plans and limits move, and note that free tiers often give you a different model from the paid tier you would actually use, so test the tier you will pay for.

The trap

The trap is the leaky comparison. Typical leaks: giving one model a better prompt than the others, judging after you have seen the brand names, testing only easy tasks, letting a single standout answer decide, and treating the result as permanent. Models are updated and renamed regularly, so a card you filled in last spring describes last spring.

What to skip

Guardrails

Sources

All 751 AI guides · JulieMango plans from £17/mo