AI Guides › Playbooks

The Income Claim Ledger: Grading "Ways to Earn With AI" Before You Spend a Weekend on Them

By Nigel Guy · 7 min read

Most "five ways to make money with AI" lists are a screenshot, a confident headline and a gap where the evidence should be. You read them, pick the one that sounds easiest, and discover three weeks later that nobody ever showed you the invoice. The list feels useful because it is specific. Specific is not the same as sourced.

The rule: never act on an income claim until you have graded what kind of evidence sits behind it, and only build on the claims that survive the grade.

This guide gives you that grading tool, then applies it to five patterns that are genuinely supported by published material. One finding up front: the published sources below describe what people do with Claude and how well it works. None of them publishes what anyone earned. Treat that gap as part of the answer.

The Ledger: four grades for any income claim

Write every claim you meet into a four-row ledger. Each claim gets one grade.

Grade What it means Example of what it looks like How much weight to give it
A: Measured A published dataset or study, with method and date An Anthropic research report with stated data period Build on it, but it describes usage, not your profit
B: Documented trial A named organisation ran it and wrote up successes and failures A published experiment with a stated outcome Learn the mechanism, not the headline
C: Named claim A named person says they earned X, no records An interview or social post Interesting; verify before copying
D: Anonymous or vague "People are making..." with no name or date Most listicles Ignore

Then add a fifth column called What would I need to see? For a C or D claim, the answer is usually an invoice, a client reference or a date range. If you cannot name it, the claim stays out of your plan.

The five patterns that pass

These are the patterns the evidence supports. Each is a mechanism, not a promise of income.

1. Sell the checkable work, not the clever work

Grade A. Anthropic's Economic Index report of 15 January 2026 (data from 13 to 20 November 2025) found computer and mathematical tasks make up roughly a third of Claude.ai usage, with software debugging alone about 6% of conversations. The same report found the top ten tasks account for 24% of conversations. The practical reading: demand clusters around a few repeatable jobs, and one of them is fixing things that can be tested.

What to do: choose work where you can tell in minutes whether the output is right (a script that runs, a spreadsheet that reconciles, a page that renders). Checkable work is where a mistake costs you least.

2. Keep a human in the loop and price your judgement

Grade A. The same report found augmented use (you and Claude working back and forth) at 52% of Claude.ai conversations by November 2025, up from 45% in August, while fully directive use fell from 39% to 32%. Anthropic suggests product changes such as file creation, memory and Skills nudged people toward collaboration. Enterprise API use, by contrast, was heavily automated.

What to do: if you are a freelancer or small business, your product is the reviewed result, not the raw output. The review is the part clients cannot easily do themselves.

3. Keep jobs short

Grade A. The report puts success at roughly 60% for tasks that take a person under an hour, falling to about 45% for tasks taking five or more human hours. It also notes complex tasks save more time when they work, which is exactly why they are tempting.

What to do: break large jobs into pieces of an hour or less, check each piece, then assemble. A long task fails quietly; a short one fails where you can see it.

4. Write the procedure down before you hand over the keys

Grade B. Anthropic's Project Vend had Claude (Sonnet 3.7) run an office shop for about a month, and the first phase lost money: it was talked into discounts, sold below cost, and invented a payment account. Anthropic said it would not hire Claudius on that showing. In phase two, with newer models, a CRM, price-research tools and required checking procedures, negative-profit weeks were largely eliminated and the shop grew to three locations. Anthropic's own lesson was that "bureaucracy matters", and that models stayed manipulable.

What to do: if you want AI to do something with money attached, give it a written checklist (verify cost, verify price, confirm before committing) and keep a person on refunds, discounts and payments.

5. Do the arithmetic on adoption, not on hope

Grade A. The report adjusts estimated productivity gains from about 1.8 to roughly 1.0 percentage points a year once task success rates are counted. For you, the equivalent is simple: count the time you spend checking and fixing before you count the time saved.

What to do: time three real jobs end to end, including review, before quoting a price.

Worked example (hypothetical)

Suppose you run a two-person bookkeeping practice and see a post claiming "AI lets you triple your client count". Ledger entry: Claim: triple clients. Grade: D, no name or date. Needed: a client count before and after, with dates. Verdict: out.

Now suppose you consider offering monthly reconciliation summaries. Ledger entry: Claim: checkable, short tasks suit Claude. Grade: A, the January 2026 Economic Index. Plan: pilot with two friendly clients, one hour per summary, you review every line, and you time the review. That plan rests on evidence, and it tells you your own numbers within a month.

A prompt to run the ledger for you

Paste a claim you have seen and let Claude grade it. Fill in [CLAIM], [WHERE_YOU_SAW_IT] and [YOUR_SKILLS].

You are a sceptical research assistant helping a UK sole trader decide whether to act on an income claim.

Claim: [CLAIM]
Where I saw it: [WHERE_YOU_SAW_IT]
My relevant skills: [YOUR_SKILLS]

Goal: tell me whether this claim is worth a pilot, and what the pilot should look like.

Steps:
1. Restate the claim in one sentence and list every number in it.
2. Grade it A (published dataset or study with method and date), B (named organisation's documented trial), C (named person, no records) or D (anonymous or vague). Explain the grade in two sentences.
3. List what evidence I would need to see to move it up one grade.
4. Identify whether the work is checkable within minutes and whether it breaks into jobs of under an hour.
5. Propose a one-month pilot with a measurable pass mark that I set before starting.

Rules:
- Do not invent figures, sources, quotes or examples. If you cannot verify something, say "unverified".
- If a detail you need is missing, ask me before answering.
- Do not promise income.

Before you reply, check that every statement is either supported by something I gave you or marked unverified. Format: a short table for the grade, then numbered lists for steps 3 to 5.

Cost, at time of writing

Claude's Free plan costs nothing. Pro is listed at $20 a month ($17 a month billed annually), about £16 a month at time of writing, and Max starts at $100 a month, about £75 at time of writing. Anthropic's page shows prices in USD excluding tax, so check the £ price at checkout. Start on Free or Pro; there is no case for Max before you have a paying pilot.

What to skip

Guardrails

Sources

All 751 AI guides · JulieMango plans from £17/mo