AI Guides › Playbooks
By Nigel Guy · 7 min read
Most people who say AI is overrated are not wrong about what they saw. They are reporting on a test they ran once, a while ago, on a free plan, with a task picked to see if it would fail. The verdict then hardens into a fact, and nobody reruns the test. It feels like healthy scepticism; in practice it is a stale benchmark you keep quoting.
The rule: an opinion about an AI tool expires. Retest it on a real task of yours, with a pass mark you set before you start, and only then decide.
This guide gives you the mechanism for that, the Recheck Card, and the five areas where the products have changed enough to be worth a rerun.
These are described from the vendors' own help pages as of 4 October 2026.
| Area | What you probably remember | What the product does now | Where to find it |
|---|---|---|---|
| Agents | A chatbot that tells you how to do the task | Tools that carry out multi-step work: browsing, filling forms, producing files | ChatGPT: tools menu, or type /agent in the box. Claude: choose Cowork from the dropdown in the message box |
| Images | Mangled text, no way to fix one bit | Edit an existing image and change a selected region | ChatGPT: More then Images, or ask in a chat; use Select in the editor |
| Voice | Dictation with a robotic reply | Full spoken conversation that can use web search and connected tools | Claude: the sound-wave icon in the chat box (all plans) |
| Desktop apps | A web page in a frame | Apps that reach local files, the browser and other programs | Claude desktop (macOS, Windows, Linux beta); the new ChatGPT desktop app combining Chat, Work and Codex (macOS and Windows) |
| Prompt ceremony | Long "magic" templates in capital letters | Plain, explicit instructions with context work better | Anthropic's prompting best-practice docs |
A few details that matter more than the headline:
What you'll pay: the free tiers cover voice, a desktop app and some image generation; agents need a paid plan. The entry paid plans from both vendors sit at roughly £20 a month at time of writing, but the pricing pages we checked showed US dollars, so confirm the pound price, including VAT, at checkout before you subscribe.
One card per area. Fill in the first three lines before you open the tool. That way you can't move the pass mark afterwards.
| Line | What you write |
|---|---|
| 1. Old verdict | What you believed and roughly when you last tested it |
| 2. Real task | A task from your actual week, not a trick question |
| 3. Pass mark | What "good enough to keep using" looks like, in observable terms |
| 4. Time box | 20–30 minutes, including setup |
| 5. Result | Pass, fail or partial, against line 3 only |
| 6. Decision | Adopt for this task, retest in six months, or drop |
How to run it:
Use this in any chatbot to turn a vague hunch into a fair test. Fill in the bracketed fields.
You are helping me run a fair retest of an AI feature I previously wrote off. I want an honest result, not encouragement.
My old verdict: [WHAT_I_BELIEVED_AND_WHEN]
The feature I'm retesting: [FEATURE, e.g. ChatGPT agent mode / Claude Cowork / voice mode]
A real task from my last two weeks: [TASK_DESCRIPTION]
How long that task usually takes me by hand: [MINUTES]
Do this in order:
1. If any field above is empty or too vague to test, ask me about it and stop. Do not guess.
2. Rewrite my task as a single clear instruction I could give the feature, with the context and the reason it matters.
3. Propose a pass mark I can check by observation (time saved, number of corrections, accuracy against a source I name). Keep it to two or three conditions.
4. List what I must check by hand before trusting the output (facts, figures, names, anything sent or deleted).
5. Name one way this test could be unfair to the tool and one way it could be unfair to me.
Format: five short numbered sections matching the steps. No preamble.
Before you answer, check: is the pass mark something I can verify without your help? If not, fix it.
Fill in: your old verdict, the feature, a real recent task and how long it normally takes.
Priya runs a three-person bookkeeping practice. Her old verdict, from about eighteen months ago, is "AI can't do anything with real files". Her card for Cowork:
| Day | Card |
|---|---|
| Monday | Prompting: rerun one old task with a plain, explained instruction instead of your saved template |
| Tuesday | Voice: talk through a draft or plan on a walk instead of typing it |
| Wednesday | Desktop app: install it and try one task that needs a local file |
| Thursday | Images: edit an existing image rather than generating from scratch |
| Friday | Agents: one bounded task, read-only, on a paid plan if you have one |