AI Guides › Prompting
Retiring Your Favourite Prompt When The Model Changes Underneath It
By Nigel Guy · 2 min read
Most people find a prompt that works well and then keep using the exact
same wording indefinitely, treating it as a fixed asset rather than
something tuned to a particular model at a particular moment. That's a
reasonable habit right up until the model underneath it changes — a new
version, a provider switch, a quiet update you weren't told about — and the
same wording starts producing something subtly different. The prompt
doesn't announce that it's gone stale. It just quietly stops working as
well as it used to.
The rule: a prompt is tuned against a specific model's habits, and when
the model changes, the prompt needs re-testing — not blind continued
trust.
The mechanism: the re-certification check
- Note which model a favourite prompt was tuned against, if you can.
Exact model names and version numbers change fast enough, and are
labelled inconsistently enough across products, that this is worth
double-checking rather than trusting memory — check the current
documentation for what you're actually running before you assume.
- When you know or suspect the underlying model has changed, rerun the
prompt on a case where you already know what a good answer looks like.
A past real example works better here than a fresh one, because you have
something concrete to compare against.
- Compare the new output against that known-good benchmark, not just
against "does this look fine on its own." Wording that once
compensated for an older model's specific quirk may now be redundant, or
actively working against a model that no longer has that quirk.
- If the output has genuinely drifted, retune the prompt as if it were
new — don't assume a single word swap restores it. Treat it as a fresh
round of testing, not a patch.
- Retire any language that was clearly a workaround for a limitation the
current model doesn't seem to have any more. Carrying old workarounds
forward is harmless at best and actively confusing at worst.
What to skip
Skip re-testing every prompt after every minor update — reserve this for
prompts that actually matter: recurring reports, anything client-facing,
anything used often enough that drift would compound. Skip assuming a
favourite prompt is broken just because its output looks different from
before — different isn't automatically worse. Check it against the
benchmark before you start retuning something that wasn't actually broken.
Guardrails
- Exactly when and how often the underlying model changes isn't something
to state as settled fact here — it varies by provider and product, and
changes often enough that it's worth checking current release notes
rather than assuming a fixed schedule.
- You won't always be told when a model has changed underneath a product
you use day to day. A periodic scheduled check on your important prompts
is more reliable than waiting to notice something feels off.
- This is about catching output-quality drift in prompts you rely on. It's
not a substitute for actually reading a provider's release notes when
they publish one — those will tell you more precisely what changed than
a benchmark comparison can infer on its own.
All 751 AI guides · JulieMango plans from £17/mo