AI Guides › Judgement & Guardrails
By Nigel Guy · 2 min read
"This tool is usually reliable" and "this specific answer is correct" get treated as the same claim, and they aren't. Trust in a tool builds up over weeks of generally good experience with it. Trust in a single output has to be earned fresh, every time, because the tool being generally good doesn't tell you anything about this particular answer, on this particular question, right now.
The rule: trust in a tool is a slow-moving average built from many uses; trust in any one output is a separate judgement made fresh each time — a good average doesn't earn a pass for the next individual answer.
| Tool trust | Output trust | |
|---|---|---|
| What it's based on | Your accumulated experience across many uses, over time | The specific claim in front of you, right now |
| How it changes | Slowly — one bad answer barely moves it, one good run barely moves it either | Instantly — this answer is either checked or it isn't, regardless of history |
| What it licenses | A general willingness to keep using the tool for this kind of task | Nothing on its own — it doesn't licence skipping the check on this specific output |
| Where it goes wrong | Overreacting to a single failure and abandoning a genuinely useful tool | Underreacting because the tool's general reputation is good, and skipping the check this time |
The failure mode worth naming specifically: high tool trust quietly converts into skipped output checks. The better a tool's track record with you, the more tempting it is to wave through its next specific answer unchecked — which is exactly backwards, since a tool you trust more is one you're also more likely to be using for something that matters.
Skip re-litigating your overall trust in a tool every single time it's wrong about something small — a single miss doesn't overturn a track record built from genuinely wide use, any more than a single hit builds one. And skip assuming that a tool you distrust generally is therefore wrong about everything specific it produces; a low-trust tool can still happen to be right about a given output, which is why the individual check matters regardless of which direction your overall trust runs.