AI Guides › Playbooks

The Six-Month Recheck: Retest Five AI Features Before You Write Them Off

By Nigel Guy · 7 min read

Most people who say AI is overrated are not wrong about what they saw. They are reporting on a test they ran once, a while ago, on a free plan, with a task picked to see if it would fail. The verdict then hardens into a fact, and nobody reruns the test. It feels like healthy scepticism; in practice it is a stale benchmark you keep quoting.

The rule: an opinion about an AI tool expires. Retest it on a real task of yours, with a pass mark you set before you start, and only then decide.

This guide gives you the mechanism for that, the Recheck Card, and the five areas where the products have changed enough to be worth a rerun.

The five things that changed

These are described from the vendors' own help pages as of 4 October 2026.

Area What you probably remember What the product does now Where to find it
Agents A chatbot that tells you how to do the task Tools that carry out multi-step work: browsing, filling forms, producing files ChatGPT: tools menu, or type /agent in the box. Claude: choose Cowork from the dropdown in the message box
Images Mangled text, no way to fix one bit Edit an existing image and change a selected region ChatGPT: More then Images, or ask in a chat; use Select in the editor
Voice Dictation with a robotic reply Full spoken conversation that can use web search and connected tools Claude: the sound-wave icon in the chat box (all plans)
Desktop apps A web page in a frame Apps that reach local files, the browser and other programs Claude desktop (macOS, Windows, Linux beta); the new ChatGPT desktop app combining Chat, Work and Codex (macOS and Windows)
Prompt ceremony Long "magic" templates in capital letters Plain, explicit instructions with context work better Anthropic's prompting best-practice docs

A few details that matter more than the headline:

What you'll pay: the free tiers cover voice, a desktop app and some image generation; agents need a paid plan. The entry paid plans from both vendors sit at roughly £20 a month at time of writing, but the pricing pages we checked showed US dollars, so confirm the pound price, including VAT, at checkout before you subscribe.

The Recheck Card

One card per area. Fill in the first three lines before you open the tool. That way you can't move the pass mark afterwards.

Line What you write
1. Old verdict What you believed and roughly when you last tested it
2. Real task A task from your actual week, not a trick question
3. Pass mark What "good enough to keep using" looks like, in observable terms
4. Time box 20–30 minutes, including setup
5. Result Pass, fail or partial, against line 3 only
6. Decision Adopt for this task, retest in six months, or drop

How to run it:

  1. Pick the task from your calendar, not your imagination. Something you did by hand in the last fortnight.
  2. Write the pass mark as a check you could hand to someone else. "Saves me at least 15 minutes and needs no more than two corrections" beats "is it useful?".
  3. Use the tool as it now asks to be used. For an agent, describe the outcome and let it plan. For prompting, give context and the reason, not ceremony.
  4. Stop at the time box. A tool that needs an hour of fiddling to pass a 20-minute task has failed this round.
  5. Record the result against line 3 and nothing else. Being impressed doesn't count, and neither does being annoyed by the interface.
  6. Date the card. A drop decision expires too. Put a reminder in six months.

A prompt to set up each recheck

Use this in any chatbot to turn a vague hunch into a fair test. Fill in the bracketed fields.

You are helping me run a fair retest of an AI feature I previously wrote off. I want an honest result, not encouragement.

My old verdict: [WHAT_I_BELIEVED_AND_WHEN]
The feature I'm retesting: [FEATURE, e.g. ChatGPT agent mode / Claude Cowork / voice mode]
A real task from my last two weeks: [TASK_DESCRIPTION]
How long that task usually takes me by hand: [MINUTES]

Do this in order:
1. If any field above is empty or too vague to test, ask me about it and stop. Do not guess.
2. Rewrite my task as a single clear instruction I could give the feature, with the context and the reason it matters.
3. Propose a pass mark I can check by observation (time saved, number of corrections, accuracy against a source I name). Keep it to two or three conditions.
4. List what I must check by hand before trusting the output (facts, figures, names, anything sent or deleted).
5. Name one way this test could be unfair to the tool and one way it could be unfair to me.

Format: five short numbered sections matching the steps. No preamble.

Before you answer, check: is the pass mark something I can verify without your help? If not, fix it.

Fill in: your old verdict, the feature, a real recent task and how long it normally takes.

Worked example (hypothetical)

Priya runs a three-person bookkeeping practice. Her old verdict, from about eighteen months ago, is "AI can't do anything with real files". Her card for Cowork:

What to do this week

Day Card
Monday Prompting: rerun one old task with a plain, explained instruction instead of your saved template
Tuesday Voice: talk through a draft or plan on a walk instead of typing it
Wednesday Desktop app: install it and try one task that needs a local file
Thursday Images: edit an existing image rather than generating from scratch
Friday Agents: one bounded task, read-only, on a paid plan if you have one

What to skip

Guardrails

Sources

All 751 AI guides · JulieMango plans from £17/mo