AI Guides › Trend Watch
By Nigel Guy · 3 min read
A tool you rely on ships an update, and the release notes call it an improvement. Sometimes your experience agrees. Sometimes the things you depended on are slower, missing, moved behind a higher plan, or just behave differently — and it's hard to tell whether the tool got worse or you're imagining it. Most people settle this by either trusting the release notes or trusting a vague feeling. Neither is evidence.
The rule: an update is judged on your tasks, not on the release notes — keep a small, fixed set of your own test tasks, rerun them after a change, and call it a downgrade only when those results say so.
Set this up once, while things are working, and reuse it after every notable update.
After an update, rerun the tasks with the same inputs and compare. AI outputs vary from run to run anyway, so run each task more than once before concluding anything has changed.
| What changed | Usually | What to check |
|---|---|---|
| New capability, old ones intact | Update | Your reference tasks still behave as before |
| Output style different, quality similar | A change, not a downgrade | Whether a prompt tweak or setting brings it back |
| Feature removed or moved to a higher plan | Downgrade for you | Whether a setting or workaround exists |
| Less reliable on your tasks | Downgrade for you | Several reruns, with saved examples |
| Faster but weaker on your tasks | Trade-off | Which matters more for this task |
| Interface moved, same results | Adjustment cost | Give it a week before judging |
"For you" matters throughout. An update can be genuinely better overall and still worse for your particular use.
Skip judging on day one; novelty and irritation both distort. Skip "it's got so much worse" threads as evidence — you can't see anyone's inputs. And skip switching tools on the strength of a single bad output.