AI Guides › Workbench
By Nigel Guy · 7 min read
Most people guard against one kind of AI failure: the made-up fact. They check a citation, catch an invented statistic, and feel protected. Meanwhile the assistant has told them their business plan is "a strong foundation", dropped a caveat that would have changed their decision, and folded the moment they said "are you sure?". None of that looks like lying, which is why it works.
The rule: name the failure modes in the prompt, give the model explicit permission to commit them less, and make it show its working so you can check.
This is our own working taxonomy, not an official one. It is built on two things that are documented: Anthropic's own research on sycophancy, which defines it as responses that match what a user believes over what is true and links it to training on human preference judgements, and Anthropic's guidance on reducing hallucinations. The three-way split is a practical convenience.
| Lie | What it sounds like | What it costs you |
|---|---|---|
| Flattery | "Great question", "this is a strong draft", agreeing with your premise | You ship the weak thing |
| Invention | A confident figure, quote, citation or feature that does not exist | You repeat a falsehood under your own name |
| Smoothing | Caveats dropped, a different question answered, a position abandoned when you push back | You decide on a tidier picture than the real one |
People catch invention because it is checkable. Flattery and smoothing are about tone and omission, so nothing in the answer is false enough to flag.
| Item | What it does | Cost at time of writing | Best for | Catch |
|---|---|---|---|---|
| The full block | Targets all three lies in one paste | Free; works on any plan, including Claude's Free plan (£0). Pro is $20 a month, about £15 at time of writing, so check the £ price at checkout | Plans, drafts, decisions, anything you will act on | Adds length to every answer |
| The one-liner | A compressed version for quick questions | Free | Everyday queries | Weaker; the model has less to hold on to |
| The challenge follow-up | Forces a second pass against its own answer | Free | After any answer you are about to rely on | Costs one extra turn |
| The claim audit | Splits an answer into checkable claims | Free | Anything with facts, figures or sources | Only as good as your checking afterwards |
Paste this at the end of any prompt, after your actual request. Nothing needs filling in, but if you add a short line about what you will use the answer for, the result improves.
Before you answer, apply these rules to everything you write.
1. Agreement is not a goal. Assess my idea, draft or assumption on its merits. If I am wrong or the work is weak, say so plainly and say why. Do not open with praise, and only praise something when you can name the specific thing that works.
2. Separate what you know from what you are guessing. Mark each important claim as one of: established, likely, or unsure. If you do not know something, say "I don't know" instead of filling the gap. Never invent names, figures, quotes, links, citations or product features. If I haven't given you a document, tell me what you are relying on memory for.
3. Surface what I might not want to hear. Name the strongest objection to my position, the biggest risk, and anything important my question left out. Do not drop caveats to make the answer tidier.
4. Answer the question I asked. If you are reframing it, tell me you are doing so and why.
5. If I push back without giving new evidence or a new argument, do not change your answer. Explain what would change it. If I give a good reason, update and say what changed your mind.
6. If you need information I haven't supplied, ask me before answering rather than assuming.
Format: a direct answer first, then "Strongest objection", then "What I'm unsure about", then "What to check myself". Keep each part short.
Self-check before sending: have you praised without evidence, stated anything you cannot support, or left out a caveat that matters? Fix it first.
Why each rule earns its place:
For quick questions, where the full block is overkill:
Be critical, not agreeable: give your honest assessment, say "I don't know" where you don't, never invent facts or sources, name the strongest objection and anything I've left out, and don't change your answer unless I give you a reason.
Use it when the cost of a wrong answer is a mildly wasted afternoon. Use the full block when the answer feeds a decision, a client or a published page.
Send this after an answer you are about to rely on. It works because a second pass with a different job catches things the first pass smoothed over.
Now argue against your own answer. List the three most likely ways it is wrong or incomplete, what evidence would show each one, and whether your confidence should be lower than you stated. Then give a revised answer, or tell me nothing changes and why.
When you have supplied a document, or the answer contains facts, turn it into something checkable. This follows Anthropic's advice to ground answers in direct quotes and retract claims that have no support.
List every factual claim in your last answer as a numbered list. For each one, give the exact supporting quote from [DOCUMENT_OR_SOURCE_I_PROVIDED], or write "no source: from memory" or "no source: inference". Then delete any claim you cannot support and show me what is left. Do not add new claims.
Fill in the name of the document or source you pasted. If you pasted nothing, expect most lines to say "from memory", which is the useful finding.