AI Guides › Playbooks
By Nigel Guy · 7 min read
Most "five ways to make money with AI" lists are a screenshot, a confident headline and a gap where the evidence should be. You read them, pick the one that sounds easiest, and discover three weeks later that nobody ever showed you the invoice. The list feels useful because it is specific. Specific is not the same as sourced.
The rule: never act on an income claim until you have graded what kind of evidence sits behind it, and only build on the claims that survive the grade.
This guide gives you that grading tool, then applies it to five patterns that are genuinely supported by published material. One finding up front: the published sources below describe what people do with Claude and how well it works. None of them publishes what anyone earned. Treat that gap as part of the answer.
Write every claim you meet into a four-row ledger. Each claim gets one grade.
| Grade | What it means | Example of what it looks like | How much weight to give it |
|---|---|---|---|
| A: Measured | A published dataset or study, with method and date | An Anthropic research report with stated data period | Build on it, but it describes usage, not your profit |
| B: Documented trial | A named organisation ran it and wrote up successes and failures | A published experiment with a stated outcome | Learn the mechanism, not the headline |
| C: Named claim | A named person says they earned X, no records | An interview or social post | Interesting; verify before copying |
| D: Anonymous or vague | "People are making..." with no name or date | Most listicles | Ignore |
Then add a fifth column called What would I need to see? For a C or D claim, the answer is usually an invoice, a client reference or a date range. If you cannot name it, the claim stays out of your plan.
These are the patterns the evidence supports. Each is a mechanism, not a promise of income.
Grade A. Anthropic's Economic Index report of 15 January 2026 (data from 13 to 20 November 2025) found computer and mathematical tasks make up roughly a third of Claude.ai usage, with software debugging alone about 6% of conversations. The same report found the top ten tasks account for 24% of conversations. The practical reading: demand clusters around a few repeatable jobs, and one of them is fixing things that can be tested.
What to do: choose work where you can tell in minutes whether the output is right (a script that runs, a spreadsheet that reconciles, a page that renders). Checkable work is where a mistake costs you least.
Grade A. The same report found augmented use (you and Claude working back and forth) at 52% of Claude.ai conversations by November 2025, up from 45% in August, while fully directive use fell from 39% to 32%. Anthropic suggests product changes such as file creation, memory and Skills nudged people toward collaboration. Enterprise API use, by contrast, was heavily automated.
What to do: if you are a freelancer or small business, your product is the reviewed result, not the raw output. The review is the part clients cannot easily do themselves.
Grade A. The report puts success at roughly 60% for tasks that take a person under an hour, falling to about 45% for tasks taking five or more human hours. It also notes complex tasks save more time when they work, which is exactly why they are tempting.
What to do: break large jobs into pieces of an hour or less, check each piece, then assemble. A long task fails quietly; a short one fails where you can see it.
Grade B. Anthropic's Project Vend had Claude (Sonnet 3.7) run an office shop for about a month, and the first phase lost money: it was talked into discounts, sold below cost, and invented a payment account. Anthropic said it would not hire Claudius on that showing. In phase two, with newer models, a CRM, price-research tools and required checking procedures, negative-profit weeks were largely eliminated and the shop grew to three locations. Anthropic's own lesson was that "bureaucracy matters", and that models stayed manipulable.
What to do: if you want AI to do something with money attached, give it a written checklist (verify cost, verify price, confirm before committing) and keep a person on refunds, discounts and payments.
Grade A. The report adjusts estimated productivity gains from about 1.8 to roughly 1.0 percentage points a year once task success rates are counted. For you, the equivalent is simple: count the time you spend checking and fixing before you count the time saved.
What to do: time three real jobs end to end, including review, before quoting a price.
Suppose you run a two-person bookkeeping practice and see a post claiming "AI lets you triple your client count". Ledger entry: Claim: triple clients. Grade: D, no name or date. Needed: a client count before and after, with dates. Verdict: out.
Now suppose you consider offering monthly reconciliation summaries. Ledger entry: Claim: checkable, short tasks suit Claude. Grade: A, the January 2026 Economic Index. Plan: pilot with two friendly clients, one hour per summary, you review every line, and you time the review. That plan rests on evidence, and it tells you your own numbers within a month.
Paste a claim you have seen and let Claude grade it. Fill in [CLAIM], [WHERE_YOU_SAW_IT] and [YOUR_SKILLS].
You are a sceptical research assistant helping a UK sole trader decide whether to act on an income claim.
Claim: [CLAIM]
Where I saw it: [WHERE_YOU_SAW_IT]
My relevant skills: [YOUR_SKILLS]
Goal: tell me whether this claim is worth a pilot, and what the pilot should look like.
Steps:
1. Restate the claim in one sentence and list every number in it.
2. Grade it A (published dataset or study with method and date), B (named organisation's documented trial), C (named person, no records) or D (anonymous or vague). Explain the grade in two sentences.
3. List what evidence I would need to see to move it up one grade.
4. Identify whether the work is checkable within minutes and whether it breaks into jobs of under an hour.
5. Propose a one-month pilot with a measurable pass mark that I set before starting.
Rules:
- Do not invent figures, sources, quotes or examples. If you cannot verify something, say "unverified".
- If a detail you need is missing, ask me before answering.
- Do not promise income.
Before you reply, check that every statement is either supported by something I gave you or marked unverified. Format: a short table for the grade, then numbered lists for steps 3 to 5.
Claude's Free plan costs nothing. Pro is listed at $20 a month ($17 a month billed annually), about £16 a month at time of writing, and Max starts at $100 a month, about £75 at time of writing. Anthropic's page shows prices in USD excluding tax, so check the £ price at checkout. Start on Free or Pro; there is no case for Max before you have a paying pilot.