AI Guides › Skills & Agents
The One Metric That Tells You A Skill Is Actually Working
By Nigel Guy · 2 min read
Most people judge whether a skill is working by how it feels to use —
smooth, fast, no obvious errors this week. That's a fine early signal and
a poor long-term one, because a skill can feel fine for months while
quietly doing the wrong thing on a class of cases nobody's happened to
check.
The rule: the one metric that actually tells you a skill is working
isn't how often it runs or how fast it feels — it's how often its output
survives contact with someone who checks it against the truth, and if
you're not measuring that, you don't actually know whether it's working.
The mechanism: the Verification Rate
- Pick a sample of the skill's outputs on a regular cadence. Not
every run — that's not sustainable and isn't the point — a defined
sample, chosen before you look at the outputs, so you're not
unconsciously picking the ones that already look fine.
- Check each one against ground truth, independently of the skill.
Whatever "correct" actually means for this task — the source document,
the real number, the person who'd know — verified by something other
than the skill's own confidence in its answer.
- Record the rate, not just the failures. "9 of 10 correct" is a
metric you can track over time. "Found one wrong one" is an anecdote
that tells you nothing about whether this month is better or worse than
last month.
- Watch the trend, not the single reading. One low sample could be
noise. A verification rate that's drifting down over several samples is
the actual signal — usually meaning the input the skill's seeing has
changed shape, even if the skill itself hasn't.
- Set a floor in advance, before you have a number to react to. Decide
what verification rate is acceptable before you see the first result —
otherwise the temptation is to retroactively decide whatever number
you got was fine.
Usage volume, run speed, and absence of visible errors are all real
signals, but they're proxies. Verification rate is the metric that
actually answers "is it working," because it's the only one measured
against reality rather than against the skill's own output.
What to skip
Skip treating "no complaints" as equivalent to a good verification rate —
silence usually means nobody checked, not that everything was right.
Skip sampling only the cases that are easy to verify; if the hard-to-check
cases are exactly the ones most likely to go wrong, a sample that avoids
them is measuring the wrong population.
Guardrails
- Verifying against ground truth takes real effort, which is exactly why
it gets skipped. There's no way around that cost if you actually want to
know the answer, rather than assume it.
- A high verification rate on old, familiar input doesn't guarantee the
same rate as the input changes — keep sampling on an ongoing basis, not
as a one-time check at launch.
- This metric tells you whether the skill's output is correct. It doesn't
tell you whether the task was worth automating in the first place —
that's a separate, earlier question.
All 751 AI guides · JulieMango plans from £17/mo