AI Guides › Playbooks
By Nigel Guy · 7 min read
Every autumn the "best AI tools of the year" lists arrive, and most people either ignore them or swap their whole setup over a weekend because one launch video looked impressive. Both feel decisive. Neither answers the narrower question: for this one job I do every week, is last year's tool still right, and what would moving cost me?
The rule: review your tools one job at a time, against what you used a year ago, and only swap a job when you can name the specific thing that changed and you have tested it on your own work.
The mechanism is the Swap Ledger: one row per job, a before column, an after column, and a verdict you have to justify in a sentence.
Posts listing which tool "replaced" which this year are claims about other people's habits, and we found no primary source that measures them, so we have not reproduced one. Even where a tool clearly improved, "better" depends on your files, budget and privacy constraints. The check below keeps working after this year's lists go stale.
One row per job you actually do, not per tool you pay for.
| Job | Before (Oct 2025) | After (Oct 2026) | What specifically changed | Switch cost | Verdict |
|---|---|---|---|---|---|
| Drafting and editing | |||||
| Research and summarising sources | |||||
| Coding or spreadsheet formulas | |||||
| Images and design | |||||
| Meetings and transcription | |||||
| Working with documents and data | |||||
| Automating repeat tasks |
| Verdict | Means | When to use it |
|---|---|---|
| Keep | Current tool stays for this job | No named change, or the test was a draw |
| Swap | Move this job to the contender | Contender won your test and switch cost is acceptable |
| Split | Use both, each for a named sub-task | Each wins on a different part of the job |
| Drop | Stop using any tool for this job | The job got faster without AI, or the output needed so much fixing it saved nothing |
A check that never returns Drop is a shopping list, not a review.
Run two or three real examples from the last month through both tools with identical instructions and files. Decide the pass mark before reading either output — "every figure matches the source", say, or "fewer than five edits". Score without labels if you can.
A chatbot can generate a shortlist, with one catch: it answers from training data with a cut-off date, so it may miss recent releases and price changes. Treat its answer as things to check, never as the verdict.
Fill in your jobs, current tools, frustrations and monthly budget in pounds:
You are helping me review which AI tools I use for which jobs. You are a sceptical adviser, not a cheerleader for any product, including the company that made you.
My jobs, each with the tool I use now and what frustrates me about it:
[JOB_1] — [CURRENT_TOOL_1] — [FRUSTRATION_1]
[JOB_2] — [CURRENT_TOOL_2] — [FRUSTRATION_2]
[ADD_MORE_ROWS]
My monthly budget for AI tools: [BUDGET_IN_GBP]
Constraints (for example client confidentiality, must work on mobile, UK data handling): [CONSTRAINTS]
What I want back:
1. For each job, say whether my current tool is still a sensible choice. Do not simply agree with me; if another tool may now fit the job better, name it and give the specific capability that would make the difference.
2. If you think I should keep what I have, say so plainly. "Keep" is a valid answer.
3. For every contender, state what I should verify on the vendor's own pricing or help pages before believing you, because your knowledge has a cut-off date and products change.
4. Never quote a price or plan limit as fact. If you mention one, label it "check current price".
Output: a table with columns Job | Current tool | Keep or test a contender | Contender (if any) | The specific reason | What to verify. Then, below the table, the single job where testing a switch is most worth my time, and why.
If any job, tool or constraint above is missing or unclear, ask me about it before you answer instead of guessing.
Before you reply, check: have you recommended a switch for any job without naming a concrete capability? Have you stated any price as certain? Fix those first.
Then design the test for the one job it flags. Fill in the job, both tools and two real examples:
You are helping me run a fair side-by-side test of two AI tools on one job.
Job: [JOB]
Tool A (current): [TOOL_A]
Tool B (contender): [TOOL_B]
Two real examples of this job I can use as test material: [EXAMPLE_1], [EXAMPLE_2]
Steps:
1. Write one set of instructions I can give both tools word for word.
2. Propose three checkable pass criteria (accuracy against source, edits needed, time to usable result) and ask me to confirm them.
3. Give me a blank scoring grid: both tools, both examples.
Do not predict a winner. Flag anything in my examples that looks confidential and suggest how to anonymise it.
Before you reply, check that the instructions in step 1 favour neither tool.
Priya runs a two-person bookkeeping practice in Leeds. This scenario is invented to show the mechanism; it is not a recommendation for any product.
| Job | Before | After | What changed | Switch cost | Verdict |
|---|---|---|---|---|---|
| Draft client emails | Chatbot A | Chatbot A | Nothing she can name | — | Keep |
| Summarise HMRC guidance | Chatbot A | Contender B | B's help pages describe citing its sources; she tested two guidance pages and B linked every claim | Saved custom instructions to recreate | Swap |
| Spreadsheet formulas | Nothing | Chatbot A | — | — | Keep (new row) |
| Meeting notes | Transcription app | Nothing | Clients asked not to be recorded | Annual plan ends in March | Drop at renewal |
One swap, one drop, and she can say exactly why.
If a row says Swap, export first.