AI Guides › Skills & Agents

The One Metric That Tells You A Skill Is Actually Working

By Nigel Guy · 2 min read

Most people judge whether a skill is working by how it feels to use — smooth, fast, no obvious errors this week. That's a fine early signal and a poor long-term one, because a skill can feel fine for months while quietly doing the wrong thing on a class of cases nobody's happened to check.

The rule: the one metric that actually tells you a skill is working isn't how often it runs or how fast it feels — it's how often its output survives contact with someone who checks it against the truth, and if you're not measuring that, you don't actually know whether it's working.

The mechanism: the Verification Rate

  1. Pick a sample of the skill's outputs on a regular cadence. Not every run — that's not sustainable and isn't the point — a defined sample, chosen before you look at the outputs, so you're not unconsciously picking the ones that already look fine.
  2. Check each one against ground truth, independently of the skill. Whatever "correct" actually means for this task — the source document, the real number, the person who'd know — verified by something other than the skill's own confidence in its answer.
  3. Record the rate, not just the failures. "9 of 10 correct" is a metric you can track over time. "Found one wrong one" is an anecdote that tells you nothing about whether this month is better or worse than last month.
  4. Watch the trend, not the single reading. One low sample could be noise. A verification rate that's drifting down over several samples is the actual signal — usually meaning the input the skill's seeing has changed shape, even if the skill itself hasn't.
  5. Set a floor in advance, before you have a number to react to. Decide what verification rate is acceptable before you see the first result — otherwise the temptation is to retroactively decide whatever number you got was fine.

Usage volume, run speed, and absence of visible errors are all real signals, but they're proxies. Verification rate is the metric that actually answers "is it working," because it's the only one measured against reality rather than against the skill's own output.

What to skip

Skip treating "no complaints" as equivalent to a good verification rate — silence usually means nobody checked, not that everything was right. Skip sampling only the cases that are easy to verify; if the hard-to-check cases are exactly the ones most likely to go wrong, a sample that avoids them is measuring the wrong population.

Guardrails

All 751 AI guides · JulieMango plans from £17/mo