AI Guides › Playbooks
By Nigel Guy · 7 min read
Most people choose an AI tool once, usually the first one that impressed them, and then spend months defending it. Every new release gets read as a threat to the decision rather than as information. It feels like loyalty and good sense; in practice it is team-shirt theatre, and the cost is quiet: you keep using the weaker tool for jobs where another one has been better for a while, and you never find out.
The rule: judge tools job by job, on your own real work, with the pass mark written down before you run the test, and put a date on every verdict.
The mechanism is the Job Scorecard, a single table you keep and revisit. It will not tell you which tool is best; that depends on your jobs and changes every few months.
One row per job. Columns:
| Column | What goes in it |
|---|---|
| Job | A specific, repeatable task ("turn call notes into a client follow-up email"), not a category ("writing") |
| Share of your AI use | Rough estimate from the last two weeks: high, medium, low |
| Test brief | The exact input you will give every tool, saved so you can reuse it |
| Pass mark | 3-4 criteria written before testing, each scored 0-2 |
| Scores | One column per tool you tested |
| Winner | The tool that earned the job, or "no clear winner" |
| Checked on | The date you ran it |
| Rematch | The date or event that triggers a retest |
Scroll back through two weeks of chat history and write down each distinct task. Be literal: summarising a tender PDF and summarising a meeting transcript are different jobs because they fail differently.
Rank the list by how often you do each job and how much a bad result costs you. Keep the top four. These should be the bulk of your real work, not the one exotic task you tried once. A tool that wins your rarest job but loses your most common one is the wrong default.
For each job, save one real input (anonymised if needed) and write the pass mark before you see any output. If you write the criteria after reading the answers, you will reward whichever answer you happened to like. Good criteria are checkable: "every figure matches the source", "under 150 words", "no invented next steps", "I could send it with fewer than three edits".
This prompt helps you draft the criteria. Fill in the job and paste your sample input.
You are helping me design a fair test for comparing AI assistants on one task I do often.
Task: [JOB_DESCRIPTION]
Who the output is for: [AUDIENCE]
What a bad result usually looks like for me: [COMMON_FAILURE]
Sample input I will use for the test: [PASTE_SAMPLE_INPUT]
Steps:
1. Restate the task in one sentence so I can confirm you understood it.
2. Propose 4 pass criteria. Each must be something I can check by reading the output, not a matter of taste. Prefer criteria tied to accuracy against the input, length, format and how much editing I would need.
3. For each criterion, describe what scores 0, 1 and 2.
4. Flag anything in my sample input that would make the test unfair (for example, it is too short to show differences, or it contains details only one tool could know from my history).
Output: a table with columns Criterion | 0 | 1 | 2, then a short list of fairness warnings.
If the task, audience or sample is missing or vague, ask me for it before writing criteria. Do not invent details about my work.
Before answering, check that every criterion is observable in the output and that none rewards a particular writing style.
Paste the identical brief into each tool on the same day. Two things keep this fair:
Copy each output into a document labelled A, B, C, strip anything that gives the tool away, and score against your pass mark. If you want a second opinion, use a scoring prompt, but treat it as a check on your own scores, not a replacement.
Fill in your criteria table, the original input, and the anonymised outputs.
You are an impartial reviewer scoring anonymised drafts against criteria I wrote before seeing them.
Original input: [PASTE_SOURCE_INPUT]
Criteria with 0/1/2 descriptions: [PASTE_CRITERIA_TABLE]
Draft A: [PASTE_OUTPUT_A]
Draft B: [PASTE_OUTPUT_B]
Draft C (optional): [PASTE_OUTPUT_C]
Steps:
1. Score each draft on each criterion, quoting the line from the draft that justifies the score.
2. For accuracy criteria, list any claim in a draft that is not supported by the original input.
3. Total the scores and say whether the gap between the top two is large enough to matter (2 points or more) or too close to call.
Output: a score table, then the unsupported claims list, then a one-line verdict.
Constraints: do not guess which product wrote which draft. Do not reward length or polish unless a criterion asks for it. If a criterion is ambiguous, say so instead of interpreting it generously.
Before answering, re-read each quoted line to confirm it supports the score you gave.
Be aware that a model scoring its own output may be lenient. Your scores count; the model's are a sense check.
Write the winner, or "no clear winner" if the gap is under two points. A close result is useful: stay with whatever is already open for that job.
A verdict without a date is how a one-off test hardens into a belief. Retest a row when any of these happens: a major model release from a tool on your list, a plan or price change, the job itself changing, or 90 days passing. The brief and criteria are saved, so a rematch is quick.
A freelance bookkeeper (invented for illustration) finds four jobs make up most of her AI use. After one afternoon of testing, her scorecard reads:
| Job | Share | Tool X | Tool Y | Winner | Checked on | Rematch |
|---|---|---|---|---|---|---|
| Client chase emails from notes | High | 7/8 | 7/8 | No clear winner | 2026-10-04 | 2027-01-04 |
| Summarise HMRC guidance pages | Medium | 5/8 | 8/8 | Tool Y | 2026-10-04 | Next major release |
| Tidy messy CSV exports | High | 8/8 | 5/8 | Tool X | 2026-10-04 | 2027-01-04 |
| First draft of engagement letters | Low | 6/8 | 6/8 | No clear winner | 2026-10-04 | 2027-01-04 |
Her conclusion is not "Tool X is better". It is "use X for CSVs, Y for guidance summaries, and either for the rest". She also notices she does not need two paid subscriptions: the free tier of Y covers her guidance volume, so she pays only for X.
Running the scorecard is free if you test on free tiers, though limits on free plans can cut a test short. At time of writing, the main individual paid tiers in the UK sit at similar levels: Google AI Pro is listed at £18.99 a month, and Claude Pro and ChatGPT Plus are each roughly £18-£20 a month depending on how VAT and currency are shown at checkout. Lower tiers also exist (ChatGPT Go, Google AI Plus) with tighter limits. Prices and tiers change often, so check the live checkout page before you pay.