AI Guides › Trend Watch
By Nigel Guy · 3 min read
When a new model is released, nearly everything you'll read is about what it does well: the benchmarks it tops, the impressive demos, the tasks it handles that the last one couldn't. You come away knowing what it's capable of at its best. What you don't come away knowing is whether it will reliably do your work — and that depends far more on where it fails than on where it shines.
The rule: a model's strengths tell you its ceiling; its weaknesses tell you whether you can depend on it. For deciding whether to use a model for real work, study the weaknesses first.
Strengths are selected for display. The launch materials, the demos, the first wave of posts — they all show the model on the tasks it was chosen to look good on. That's not dishonest, but it's a curated sample.
Weaknesses are less curated. They show up in the limitations section of the documentation, in the model card if there is one, in bug reports, and in the quieter posts from people who tried it on ordinary work. And a weakness that overlaps with your work is a direct predictor of the time you'll spend checking and fixing its output.
Build this for any new model you're considering for a real task. It takes an hour or so, and it's more useful than any leaderboard.
| Failure type | What it looks like | Why it matters |
|---|---|---|
| Confident errors | Wrong facts stated fluently | Costs you verification time on every output |
| Instruction drift | Follows the brief early, loses it later | Longer tasks become unreliable |
| Format breakage | Ignores a required structure | Breaks anything downstream that parses its output |
| Refusal or over-caution | Declines reasonable requests in your domain | Blocks legitimate work unpredictably |
| Inconsistency | Different quality on repeat runs of the same prompt | Makes it hard to build a routine |
Skip the launch-week demos as evidence for your decision — they're evidence of the ceiling, not the floor. Skip benchmark comparisons for tasks you don't do; see "Reading A Benchmark Chart Without Being Fooled By It." And skip generalising from a single failure or a single success. Run each task more than once before you conclude anything.