AI Guides › Playbooks

The Full-Cost Ledger: Reading an AI "Price Per Result" Claim

By Nigel Guy · 7 min read

"Ten long-standing maths problems solved for about $2,000" is a headline built to make you think the price of intelligence has collapsed. Most people either swallow it whole ("so I can buy breakthroughs by the dollar") or wave it away ("hype"). Both skip the useful bit: what that number actually counts, and what it leaves out.

The rule: a price per result is only the price of the final run. Before you act on one, write down every other cost that sat behind it, and cost your own task the same way.

What was actually claimed

On 1 August 2026 OpenAI published "Ten advances in mathematics and theoretical computer science". It says the results came from an internal version of Astra, which it calls its next major model, and that "the total number of tokens needed to find solutions to these problems would cost roughly $2,000 at Sol API rates" (about £1,500 at time of writing; check the current exchange rate). The ten problems span sphere packing, coding theory, group theory, operator algebras, circuit complexity, quantum games, lattice problems, convex geometry and extremal combinatorics, and three of them are Erdős problems (146, 180 and 183). Humans then prepared manuscripts with the same model, and the model produced Lean formalisations, which are machine-checkable certificates, published at OpenAI's ten-proofs GitHub repository.

Some things are not confirmed by anything I could read. Astra was unreleased when the post was written, so you cannot buy it. I could not verify what "Sol API rates" are, so the dollar figure cannot be reproduced. And the figure covers the tokens that found the solutions, not the experiments that did not work, nor the staff, nor the checking. That is a vendor's own statement about its own work. Treat it as a data point, not an audit.

For scale, the same company's earlier Navier–Stokes announcement was reported by the BBC as roughly 10,000 agents for 88 hours, which the BBC estimated at about $10m (£7.3m) at OpenAI's published pricing. Same vendor, same period, a price five thousand times higher. "AI maths is cheap" and "AI maths is dear" are both true, depending on the problem.

The Full-Cost Ledger

Use this whenever someone quotes you a cost per AI result, including your own business case. Five lines, filled in in this order.

Line Question In the maths example
1. Run cost What did the winning run cost in tokens? About $2,000 for ten problems, per OpenAI
2. Search cost How many attempts failed first, and how was the problem chosen? Not stated. OpenAI says problems were evaluated while the model was in development
3. Access cost Can you buy the same capability today, at that rate? No: unreleased model, rates not independently verifiable
4. Checking cost Who verified the result, and how long did it take? Humans plus Lean certificates; external mathematicians checked the earlier unit-distance proof
5. Framing cost Who knew which question was worth asking? Researchers selected well-studied open problems. The model did not choose them for you

If a line is blank, the claim is incomplete, not false. Write "unknown" and carry on.

Step by step

  1. Copy the claim exactly. Number, unit, what it covers ("tokens needed to find solutions"), and the source. Not the social post summarising it.
  2. Fill lines 1 to 3 from the primary source. If the source does not say, the answer is "not stated".
  3. Fill line 4 honestly. For anything you will rely on, ask who checks it and how. Proof assistants such as Lean make checking cheap for maths. Your marketing copy or contract summary has no Lean. Checking will be a person, and that person's time is your biggest line.
  4. Fill line 5. The expensive human skill is often choosing a question worth solving and recognising a good answer. That did not get cheaper just because tokens did.
  5. Re-run the ledger on your own task, with your own numbers (the prompt below does this).
  6. Decide on the total, not on line 1.

Worked example (hypothetical)

A small UK consultancy wants an AI model to review 40 past client proposals and flag pricing inconsistencies. A colleague says: "It's pennies in tokens, we should just do it."

Verdict: still worth doing, but the real budget is the reviewer's time. The "pennies" claim was true and irrelevant.

A prompt to run the ledger on your own task

Fill in the square brackets, then paste into a chat model you already pay for.

You are a careful operations analyst helping a small UK business decide whether an AI task is worth doing.

Context: I want to use AI for [TASK]. A source says AI can do this for about [QUOTED_COST] per result. The source is [SOURCE_NAME_OR_LINK]. Our hourly cost for a person who could check the output is [HOURLY_RATE_GBP]. Volume per month: [VOLUME].

Goal: produce a five-line "full-cost ledger" for this task and a clear go / pilot / no-go recommendation.

Steps:
1. Restate the quoted cost and say exactly what it does and does not appear to cover. If you are unsure, say so.
2. Build the ledger: run cost, search cost (failed or repeated attempts), access cost, checking cost, framing cost (who defines success and picks the questions).
3. For each line give an estimate in pounds per month with a low and high figure, and mark each as "from my inputs" or "my assumption".
4. Identify the single biggest line and what would shrink it.
5. Recommend go, pilot or no-go, and describe a one-week pilot with a pass threshold set before it starts.

Rules:
- If any input in square brackets is missing, ask me for it before answering. Do not guess.
- Do not invent prices, statistics or benchmarks. If you need a current price, tell me which page to check.
- Keep the answer under 400 words, with the ledger as a table.

Before you answer, check that every number is traceable to my inputs or labelled as an assumption.

What to skip

Guardrails

Sources

All 751 AI guides · JulieMango plans from £17/mo