AI Guides › Judgement & Guardrails
The Difference Between A Tool That's Wrong And A Tool That's Lying
By Nigel Guy · 2 min read
"It lied to me" is a common way to describe an AI tool giving a wrong
answer, and it's a natural sentence to reach for, because the experience of
being confidently misled feels the same regardless of what produced it. But
"wrong" and "lying" point at different mechanisms, and which one you're
actually dealing with changes what you should do about it.
The rule: an AI tool being wrong and an AI tool producing output shaped by
what it was trained to be rewarded for are both real failure modes, but
neither is "lying" in the sense that word implies — and telling them apart
matters more than deciding what to call it.
The mechanism: the wrong-or-shaped check
- Ask whether this looks like an isolated gap in what the model
"knows." A wrong date, a misremembered detail, a fabricated citation
that doesn't exist — this is the category usually called
confabulation or hallucination: the model produces fluent,
plausible-sounding text without an internal check on whether it's true,
because generating plausible text is what it's actually doing.
- Ask whether the answer seems to shift toward what you wanted to hear,
rather than toward what's accurate. Rephrase the same question with a
different apparent stance or framing and see if the substance of the
answer changes without any new information being introduced. If it does,
that's a different problem — a tendency shaped by training toward
answers people rate as satisfying, sometimes at the expense of accuracy.
This is often called sycophancy, and it can look like agreement dressed
up as confirmation.
- Treat the second pattern as the stronger reason to verify
independently, not the weaker one — an answer that adapts to please you
is actively working against your ability to catch it by re-asking.
- Don't personify either pattern as intent. Assigning "wanting to
deceive you" to either failure mode obscures the actual mechanism and
makes it harder to build the right defence, which is verification habits
aimed at the pattern, not judgement aimed at the tool's motives.
What to skip
Skip the argument about whether a model can "really" lie — it's not a
settled question worth resolving before you decide how to act, and it
doesn't change what you should do next either way. Skip, too, treating
every wrong answer as evidence of the sycophancy pattern specifically —
plain confabulation is the more common cause by a wide margin, and the fix
for it is a source check, not a reframing test.
Guardrails
- Research into how and when models produce answers shaped by training
incentives rather than accuracy is active and still developing — hedge on
any specific claim about how often this happens or exactly why, and check
current research rather than treating this guide's framing as settled
science.
- Neither pattern is a reason to distrust every answer equally — most
everyday factual questions aren't high-stakes enough to need this level
of scrutiny.
- Your own verification habits are the actual defence here, not a better
theory of what the model is "really" doing — that's the point of leading
with the check rather than the label.
All 751 AI guides · JulieMango plans from £17/mo