AI Guides › Judgement & Guardrails

Why Consensus Between Two AI Tools Isn't Proof Of Anything

By Nigel Guy · 2 min read

Running the same question through a second tool and getting the same answer back feels like independent confirmation — you asked twice, from two different places, and got one result. It's a natural thing to reach for, and it produces a very specific, very misleading kind of confidence, because the two tools were never actually independent in the way that feeling implies.

The rule: two AI tools agreeing tells you they were shaped by similar training data and similar incentives, not that the answer is correct — real independent verification comes from an outside source, not a second model.

The mechanism: the independence check

  1. Ask whether the two tools' training and sources plausibly overlap. Large language models trained on broadly similar internet-scale data, with broadly similar techniques, are not independent witnesses in the way two people who investigated separately would be. Agreement between them is closer to two students who read the same textbook giving the same answer than to two independent experiments confirming a result.
  2. Restate the reasoning behind each answer, not just the conclusion. If both tools reached the same conclusion through the same underlying assumption, and that assumption is wrong, agreement doesn't help you — it just means the same error appeared twice.
  3. Watch for shared blind spots specifically. Some errors are common precisely because they're the statistically likely-sounding answer rather than the correct one — which is exactly the kind of error two models trained the same general way are prone to making identically.
  4. Treat genuine independent verification as coming from outside the category entirely — a primary source, a domain expert, your own direct check — rather than from asking a third model the same question again.

What to skip

Skip treating three or four tools agreeing as proportionally stronger evidence than two — past a certain point you're not adding independent data, you're just sampling the same underlying pattern repeatedly and watching it recur. And skip using consensus-checking as your only verification step for anything that actually matters; it's a reasonable first filter for catching an obvious, isolated error, not a substitute for tracing a claim to a real source.

Guardrails

All 751 AI guides · JulieMango plans from £17/mo