AI Guides › Playbooks

The Leverage Ledger: Five Sums That Replace Vanity Metrics

By Nigel Guy · 6 min read

Most people measuring their AI use count what is easy to count: prompts sent, documents generated, tools subscribed to, hours "saved" by feel. Those numbers rise whether or not your business improves, which is exactly why they feel fine while telling you nothing. Worse, the feeling itself is unreliable: in a 2025 randomised trial by METR, 16 experienced developers working on 246 real tasks took 19% longer when AI tools were allowed, yet afterwards still believed AI had made them about 20% faster.

The rule: track only numbers that end in a sum you can redo from your own records, and pair every speed number with a quality number.

That last half is borrowed from software delivery. DORA's published metrics deliberately pair throughput measures (how fast changes ship) with instability measures (how often they fail or need rework), because speed on its own flatters you. The ledger below does the same for one person or a small team using AI.

The Leverage Ledger

Five lines, one sum each, reviewed monthly. Every figure in the worked example further down is a hypothetical placeholder, labelled as such. It is not anyone's real result and not a benchmark.

# Number The sum What it guards against
1 Net hours returned (old hours per task − new hours per task, including checking and fixing) × tasks per month Counting drafting time saved but not review time added
2 Rework rate outputs that needed material fixes ÷ outputs shipped Speed that quietly lowers quality
3 Cost per finished output (tool spend + your hours × your hourly cost) ÷ outputs accepted Ignoring your own time and paying for unused seats
4 Revenue per working hour revenue in the month ÷ all hours worked in the month Busy-ness that does not turn into money
5 Reinvestment rate freed hours spent on paid or compounding work ÷ hours freed Saved time that evaporates into more admin

1. Net hours returned

Time a task by the clock three times before AI and three times after, and use the average. "After" starts when you open the tool and ends when the work is ready to send, so prompting, waiting, reading the output properly and correcting it all count. If you cannot time it, you do not yet know the number; do not fill the gap with a guess.

2. Rework rate

Define "material fix" before you start: anything that changed a fact, a number, the structure or the tone enough that you would have been embarrassed to send the draft as it came. Typos do not count. Count honestly for a full month, then divide. A rising rework rate with rising net hours means you are trading quality for time, and the trade is rarely visible from the inside.

3. Cost per finished output

Add the subscription and usage costs attributable to this work, add your own hours at an honest hourly cost, and divide by outputs that were actually accepted, not generated. Price the tools from your current invoices; plan prices and limits change often, so check the £ amount on your own bill rather than a guide. Count only outputs that left the building: a draft you binned is cost with no output.

4. Revenue per working hour

Use total hours, including the admin and the "quick" prompting sessions in the evening. This is the only number on the list that can move for reasons unrelated to AI, which is the point: it keeps the other four honest. If numbers 1 to 3 improve for three months and this one is flat, the gains are not reaching the business.

5. Reinvestment rate

Of the hours number 1 says you freed, how many went into work that earns or compounds: selling, delivery for paying clients, building a reusable asset? Hours that went to inbox tidying or to tweaking prompts do not count. This is the number most people skip, and it is where leverage is either realised or wasted.

Worked example (hypothetical)

Imagine a one-person bookkeeping-and-reporting service. These figures are invented to show the arithmetic.

Reading it: the tool is genuinely helping, since cost per output fell and revenue per hour rose. But 60% of the freed time leaked away, and one in five outputs needed fixing. The next move is a better review checklist and a fixed slot for the freed hours, not another subscription.

Where to start

Start with number 1 and number 2 together, for one task you repeat at least ten times a month. That is a single afternoon of setup: a spreadsheet with one row per task and columns for minutes taken and whether it needed a material fix. Add numbers 3 to 5 once you have a month of rows. The trap is starting with number 4, because it is the one you care about and the one least connected to any single decision. You will see it move and have no idea why.

A note on the tracking itself: if you are measuring how your whole team uses AI rather than your own work, tell them what you are counting and why, and keep to output-level numbers, not conversation-level ones.

What to skip

Guardrails

Sources

All 751 AI guides · JulieMango plans from £17/mo