AI Guides › Trend Watch
The One Habit Worth Building Before The Next Big Model Release
By Nigel Guy · 2 min read
Every major model release follows the same script for most people: a wave
of posts, a few impressive screenshots, a quick try on whatever comes to
mind, and a verdict of "wow" or "meh" by the end of the day. Then nothing
changes in how you actually work, because a random prompt on launch day
tells you almost nothing about whether the new model is better at your
real tasks. It feels like keeping up. It's mostly spectating.
The rule: keep a small, fixed set of your own real tasks — a personal
test kit — and run it on every new model you're considering, so each
release gets judged against your work instead of someone else's demo.
Building the Personal Test Kit
You build this once, while nothing is launching, and reuse it every time
something does.
- Pick five to eight tasks you actually do. Not showcase tasks — the
ordinary ones. A report summary, a tricky email, a data clean-up, a
piece of code you maintain, whatever reflects your week.
- Save the exact inputs. Same document, same prompt, same files. If
the input changes between tests, the comparison is meaningless.
- Write down what "good" looks like for each, before testing. A short
note: must catch these three points, must keep this tone, must not
invent figures. Setting this in advance stops you grading on vibes.
- Keep your current best output as a baseline. Whatever your existing
tool produces now is the bar a new model has to clear.
- Include at least one task where tools usually fail you. That's
where real improvements show up — or don't.
Running it on release day (or week)
| Step |
What you do |
| Run |
Put each saved input through the new model, unchanged. |
| Score |
Check each output against your pre-written "good" note. |
| Compare |
Put it beside your baseline. Better, same, or worse? |
| Note cost and speed |
Does the improvement come with a price or delay you'd feel? |
| Decide |
Switch only if it clearly wins on tasks that matter to you. |
The whole run should take under an hour. If it takes longer, your kit is
too big.
What to skip
- Skip launch-day testing on random prompts. Asking a new model a
riddle or a party-trick question tells you about riddles.
- Skip other people's benchmark rankings as your deciding factor. They
measure general ability; your kit measures fit.
- Skip re-testing every minor update. Save the kit for releases you're
genuinely considering adopting.
Guardrails
- Your kit reflects your work today. Refresh a task or two every few
months as your work changes, or it'll start testing the wrong things.
- A small kit gives you a signal, not a statistical result. If outcomes are
close, run a couple of tasks again before deciding.
- Don't put confidential material in your kit if you'll be testing tools
whose data terms you haven't checked. Use realistic but safe versions.
- Early access and launch periods can behave differently from normal use —
if something looks unusually good or bad, recheck after a week.
All 751 AI guides · JulieMango plans from £17/mo