AI Guides › Trend Watch

What A New Model's Weaknesses Tell You That Its Strengths Don't

By Nigel Guy · 3 min read

When a new model is released, nearly everything you'll read is about what it does well: the benchmarks it tops, the impressive demos, the tasks it handles that the last one couldn't. You come away knowing what it's capable of at its best. What you don't come away knowing is whether it will reliably do your work — and that depends far more on where it fails than on where it shines.

The rule: a model's strengths tell you its ceiling; its weaknesses tell you whether you can depend on it. For deciding whether to use a model for real work, study the weaknesses first.

Why weaknesses carry more information

Strengths are selected for display. The launch materials, the demos, the first wave of posts — they all show the model on the tasks it was chosen to look good on. That's not dishonest, but it's a curated sample.

Weaknesses are less curated. They show up in the limitations section of the documentation, in the model card if there is one, in bug reports, and in the quieter posts from people who tried it on ordinary work. And a weakness that overlaps with your work is a direct predictor of the time you'll spend checking and fixing its output.

The Weakness Map

Build this for any new model you're considering for a real task. It takes an hour or so, and it's more useful than any leaderboard.

  1. Read the official limitations first. Most model providers publish something — a model card, a system card, a "known limitations" section. Note anything relevant to your work. Companies tend to be more candid here than in launch posts.
  2. List the three tasks you'd actually use it for. Real ones, from your own recent work — not a generic test.
  3. Run each task and look for these failure types:
Failure type What it looks like Why it matters
Confident errors Wrong facts stated fluently Costs you verification time on every output
Instruction drift Follows the brief early, loses it later Longer tasks become unreliable
Format breakage Ignores a required structure Breaks anything downstream that parses its output
Refusal or over-caution Declines reasonable requests in your domain Blocks legitimate work unpredictably
Inconsistency Different quality on repeat runs of the same prompt Makes it hard to build a routine
  1. Compare against what you use now. The question isn't "is this model good?" but "does it fail less, or less expensively, on my tasks than the one I have?"
  2. Write down one sentence per task: use it, use it with checking, or don't use it for this.

What to skip

Skip the launch-week demos as evidence for your decision — they're evidence of the ceiling, not the floor. Skip benchmark comparisons for tasks you don't do; see "Reading A Benchmark Chart Without Being Fooled By It." And skip generalising from a single failure or a single success. Run each task more than once before you conclude anything.

Guardrails

All 751 AI guides · JulieMango plans from £17/mo