AI Guides › Playbooks

The Token-Multiplier Ledger: Costing a Claude Model Switch

By Nigel Guy · 7 min read

Most people cost a model switch by comparing two price lists and stopping there. That is how a headline like "Sonnet 5 is going up 50%" gets treated as the whole story, when the bill also depends on how many tokens the new model counts for the same text. A price per token is only half of a cost.

The rule: never cost a model switch from the price list alone; multiply the price by a token count you measured on your own prompts.

First, a correction to the premise

The Sonnet 5 price rise you may have heard about did not happen. Anthropic's pricing page says the $2 / $10 per million input / output tokens, announced at launch as introductory pricing through 31 August 2026, is now the standard price, and that the scheduled increase to $3 / $15 on 1 September 2026 "will not occur". Anthropic's launch page dates the change to 10 August 2026. At time of writing, Sonnet 5 costs $2 per million input tokens (about £1.50) and $10 per million output tokens (about £7.50). Those £ figures are rough conversions; check the £ price at checkout or on your invoice.

So the useful question is no longer "how bad is the rise?" It is "what does a switch to a differently-tokenised model really cost me?" That question outlives this particular news item. The ledger below works for any switch.

What the published numbers say

Model Input per million tokens Output per million tokens Cache read per million
Claude Sonnet 4.6 $3 (about £2.25) $15 (about £11.25) $0.30
Claude Sonnet 5 $2 (about £1.50) $10 (about £7.50) $0.20
Claude Sonnet 5.5 $2 (about £1.50) $10 (about £7.50) $0.20

Prices are from Anthropic's pricing page, all in USD; £ figures are approximate at time of writing. The Batch API takes 50% off input and output for Sonnet 5 ($1 / $5), and the pricing page says batch and caching discounts stack.

The catch sits underneath the table. Anthropic says Claude 4.7 and later models use a newer tokenizer that produces roughly 30% more tokens for the same text, depending on content, and that Sonnet 4.6 and earlier use the old one. Anthropic's Sonnet 5 announcement describes the same input as roughly 1.0 to 1.35 times the tokens of Sonnet 4.6, depending on content type. So the sticker price is lower, but the count is higher.

The Token-Multiplier Ledger

The ledger is one row per workload, with five columns you fill from your own data.

  1. Workload. One recurring job, such as "summarise support tickets" or "draft product descriptions". Do not blend jobs; their multipliers differ.
  2. Old tokens. Input and output tokens per run on the current model, from your usage page or API usage fields.
  3. Multiplier. New-model tokens divided by old-model tokens for the same prompts. Measure it; do not borrow 1.3.
  4. Unit price ratio. New price divided by old price.
  5. Real cost ratio. Multiplier times unit price ratio. Below 1.0, you save. Above 1.0, you pay more.

How to measure the multiplier

Anthropic's token counting endpoint (/v1/messages/count_tokens) is free to use, subject to rate limits, and all active models support it. The documentation tells you to count the same request twice, once per model, and compare the two input_tokens values. It also says the count is an estimate that can differ slightly from what is billed, and that you should recount against the model you plan to use rather than reuse old counts.

Two limits to know about. The endpoint rejects server tools such as web search, the MCP connector, and image or document blocks sourced by URL or file ID (send those as base64). And it counts input only. Output tokens you must measure by running a sample through each model and reading the usage field.

The break-even line

Sonnet 4.6 to Sonnet 5 has a unit price ratio of 2/3, about 0.67 on both input and output. Multiply by your multiplier:

Measured multiplier Real cost ratio Meaning
1.00 0.67 About a third cheaper
1.30 0.87 About 13% cheaper
1.35 0.90 About 10% cheaper
1.50 1.00 Break-even
1.60 1.07 About 7% dearer

Break-even is a multiplier of 1.5. Anthropic's stated range tops out at 1.35, so on the published numbers a plain 4.6-to-5 switch should come out cheaper per task. Your own prompts decide whether you sit inside that range.

Worked example (hypothetical)

A small agency runs a nightly job that rewrites 500 product descriptions. These figures are invented for illustration.

The saving is real but smaller than the 33% the price list implies, because both columns grew. Had the old rewrite job used a lot of reasoning or long outputs, the output column would dominate, so the ledger tracks input and output separately.

Why the earlier scheduled rise would have hurt more than 50%

This is the part the headline missed, and it shows why the ledger exists. Going from $2 to $3 is a 50% rise per token. But if Sonnet 5 counts 1.3 times the tokens, the same work would have cost $3 x 1.3 = $3.90 per million input-equivalent against Sonnet 4.6's $3: about 30% more than the model it replaced, and about 95% more than the launch price. The per-token rise and the token-count rise multiply. That scenario was cancelled, but the multiplication still applies whenever a vendor changes either number.

A prompt to build your ledger

Fill in the bracketed parts and paste it into Claude along with a usage export. Do not paste customer data, only counts.

You are a careful cost analyst helping me compare two Claude models.

Context: I run [WORKLOAD_DESCRIPTION]. I am considering switching from [OLD_MODEL] to [NEW_MODEL].

Inputs I will give you: per-run token counts on the old model ([OLD_INPUT_TOKENS] input, [OLD_OUTPUT_TOKENS] output), token counts for the same prompts on the new model ([NEW_INPUT_TOKENS] input, [NEW_OUTPUT_TOKENS] output), runs per month ([RUNS_PER_MONTH]), and per-million-token prices for both models that I have checked on the official pricing page: [OLD_PRICES] and [NEW_PRICES]. Also: share of traffic using batch processing ([BATCH_SHARE]) and cache reads ([CACHE_SHARE]).

Goal: a Token-Multiplier Ledger row for this workload showing the input multiplier, the output multiplier, the unit price ratio, the real cost ratio and the monthly cost on each model.

Steps:
1. List any input I have not given you and ask me for it. Do not guess numbers.
2. Calculate the multipliers separately for input and output.
3. Calculate monthly cost on both models, showing the arithmetic.
4. State the break-even multiplier for these prices.
5. Say which side of break-even I am on and by how much.

Output: one markdown table, then three plain sentences of conclusion. Keep currency in the units I supplied; do not convert currencies unless I give you a rate.

Constraints: do not invent prices or discounts. Flag any figure that looks like a typo. If my sample is under about 20 runs, say the result is a rough estimate.

Before answering, check that every number in your table traces to something I gave you.

What to skip

Guardrails

Sources

All 751 AI guides · JulieMango plans from £17/mo