AI Guides › Playbooks

The Job Scorecard: Comparing AI Tools on Your Own Work

By Nigel Guy · 7 min read

Most people choose an AI tool once, usually the first one that impressed them, and then spend months defending it. Every new release gets read as a threat to the decision rather than as information. It feels like loyalty and good sense; in practice it is team-shirt theatre, and the cost is quiet: you keep using the weaker tool for jobs where another one has been better for a while, and you never find out.

The rule: judge tools job by job, on your own real work, with the pass mark written down before you run the test, and put a date on every verdict.

The mechanism is the Job Scorecard, a single table you keep and revisit. It will not tell you which tool is best; that depends on your jobs and changes every few months.

The Job Scorecard

One row per job. Columns:

Column What goes in it
Job A specific, repeatable task ("turn call notes into a client follow-up email"), not a category ("writing")
Share of your AI use Rough estimate from the last two weeks: high, medium, low
Test brief The exact input you will give every tool, saved so you can reuse it
Pass mark 3-4 criteria written before testing, each scored 0-2
Scores One column per tool you tested
Winner The tool that earned the job, or "no clear winner"
Checked on The date you ran it
Rematch The date or event that triggers a retest

Step 1 — List the jobs you actually do

Scroll back through two weeks of chat history and write down each distinct task. Be literal: summarising a tender PDF and summarising a meeting transcript are different jobs because they fail differently.

Step 2 — Keep the four that make up most of your use

Rank the list by how often you do each job and how much a bad result costs you. Keep the top four. These should be the bulk of your real work, not the one exotic task you tried once. A tool that wins your rarest job but loses your most common one is the wrong default.

Step 3 — Write the test brief and pass mark first

For each job, save one real input (anonymised if needed) and write the pass mark before you see any output. If you write the criteria after reading the answers, you will reward whichever answer you happened to like. Good criteria are checkable: "every figure matches the source", "under 150 words", "no invented next steps", "I could send it with fewer than three edits".

This prompt helps you draft the criteria. Fill in the job and paste your sample input.

You are helping me design a fair test for comparing AI assistants on one task I do often.

Task: [JOB_DESCRIPTION]
Who the output is for: [AUDIENCE]
What a bad result usually looks like for me: [COMMON_FAILURE]
Sample input I will use for the test: [PASTE_SAMPLE_INPUT]

Steps:
1. Restate the task in one sentence so I can confirm you understood it.
2. Propose 4 pass criteria. Each must be something I can check by reading the output, not a matter of taste. Prefer criteria tied to accuracy against the input, length, format and how much editing I would need.
3. For each criterion, describe what scores 0, 1 and 2.
4. Flag anything in my sample input that would make the test unfair (for example, it is too short to show differences, or it contains details only one tool could know from my history).

Output: a table with columns Criterion | 0 | 1 | 2, then a short list of fairness warnings.

If the task, audience or sample is missing or vague, ask me for it before writing criteria. Do not invent details about my work.
Before answering, check that every criterion is observable in the output and that none rewards a particular writing style.

Step 4 — Run the same brief through each tool, cleanly

Paste the identical brief into each tool on the same day. Two things keep this fair:

Step 5 — Score without knowing which is which

Copy each output into a document labelled A, B, C, strip anything that gives the tool away, and score against your pass mark. If you want a second opinion, use a scoring prompt, but treat it as a check on your own scores, not a replacement.

Fill in your criteria table, the original input, and the anonymised outputs.

You are an impartial reviewer scoring anonymised drafts against criteria I wrote before seeing them.

Original input: [PASTE_SOURCE_INPUT]
Criteria with 0/1/2 descriptions: [PASTE_CRITERIA_TABLE]
Draft A: [PASTE_OUTPUT_A]
Draft B: [PASTE_OUTPUT_B]
Draft C (optional): [PASTE_OUTPUT_C]

Steps:
1. Score each draft on each criterion, quoting the line from the draft that justifies the score.
2. For accuracy criteria, list any claim in a draft that is not supported by the original input.
3. Total the scores and say whether the gap between the top two is large enough to matter (2 points or more) or too close to call.

Output: a score table, then the unsupported claims list, then a one-line verdict.

Constraints: do not guess which product wrote which draft. Do not reward length or polish unless a criterion asks for it. If a criterion is ambiguous, say so instead of interpreting it generously.
Before answering, re-read each quoted line to confirm it supports the score you gave.

Be aware that a model scoring its own output may be lenient. Your scores count; the model's are a sense check.

Step 6 — Record the winner and the date

Write the winner, or "no clear winner" if the gap is under two points. A close result is useful: stay with whatever is already open for that job.

Step 7 — Set rematch triggers

A verdict without a date is how a one-off test hardens into a belief. Retest a row when any of these happens: a major model release from a tool on your list, a plan or price change, the job itself changing, or 90 days passing. The brief and criteria are saved, so a rematch is quick.

Worked example (hypothetical)

A freelance bookkeeper (invented for illustration) finds four jobs make up most of her AI use. After one afternoon of testing, her scorecard reads:

Job Share Tool X Tool Y Winner Checked on Rematch
Client chase emails from notes High 7/8 7/8 No clear winner 2026-10-04 2027-01-04
Summarise HMRC guidance pages Medium 5/8 8/8 Tool Y 2026-10-04 Next major release
Tidy messy CSV exports High 8/8 5/8 Tool X 2026-10-04 2027-01-04
First draft of engagement letters Low 6/8 6/8 No clear winner 2026-10-04 2027-01-04

Her conclusion is not "Tool X is better". It is "use X for CSVs, Y for guidance summaries, and either for the rest". She also notices she does not need two paid subscriptions: the free tier of Y covers her guidance volume, so she pays only for X.

What it costs

Running the scorecard is free if you test on free tiers, though limits on free plans can cut a test short. At time of writing, the main individual paid tiers in the UK sit at similar levels: Google AI Pro is listed at £18.99 a month, and Claude Pro and ChatGPT Plus are each roughly £18-£20 a month depending on how VAT and currency are shown at checkout. Lower tiers also exist (ChatGPT Go, Google AI Plus) with tighter limits. Prices and tiers change often, so check the live checkout page before you pay.

What to skip

Guardrails

Sources

All 751 AI guides · JulieMango plans from £17/mo