AI Guides › Getting Started

Setting Expectations: What AI Gets Right On The First Try, And What It Never Will

By Nigel Guy · 2 min read

New users tend to swing between two extremes: treating the first answer as gospel because it sounded confident and fluent, or writing the whole thing off after one answer turned out wrong. Both reactions skip the actual work, which is learning what kind of task tends to land well on the first try and what kind never will, however well you ask.

The rule: calibrate by task type, not by how confident the answer sounds — confidence and correctness are not the same signal here, and the sooner you stop reading tone as a proxy for accuracy, the better your judgement gets.

What tends to land well first try

Usually reliable first-pass Usually needs checking, always
Rewriting, summarising, or restructuring text you provide Any specific fact, date, statistic, or quote it generates rather than you providing
Explaining a well-established concept Anything about very recent events or fast-moving specifics (pricing, current features, current model names)
Drafting something you'll edit anyway Numbers in a calculation-heavy task — recheck the arithmetic, don't just trust the working
Brainstorming options, angles, or structures Anything where it's citing a source, name, or reference you can't independently verify
Coding help you can run and test Claims about what a specific person said, did, or believes, unless you gave it that information

The mechanism: calibrate, don't just react

  1. Separate "did it sound right" from "is it right." Fluent, confident phrasing is what the model is good at producing regardless of accuracy — treat it as neutral, not as evidence.
  2. Match your checking effort to the cost of being wrong, not to how the answer felt. A wrong fact in a low-stakes brainstorm costs you nothing; the same wrong fact in something you send to a client costs you real trust.
  3. Track your own pattern, not just the model's. After a few weeks, you'll know which of your typical tasks it handles well first try and which ones you always need to correct — that's the actually useful calibration, more specific than any general table.
  4. Recalibrate periodically. Models and features change, and what needed heavy checking six months ago may not now, and vice versa — don't let an old impression harden into a permanent rule.

What to skip

Don't assume a task category is safe just because it's on the "usually reliable" side of the table above — that's a starting steer, not a guarantee, and it doesn't excuse skipping the check on something that matters. Don't over-correct into distrusting everything either; that just means you're doing the verification work yourself with none of the drafting help.

Guardrails

All 751 AI guides · JulieMango plans from £17/mo