AI Guides › Playbooks
By Nigel Guy · 6 min read
Most people who hit their Claude limit assume they did too much work. Usually they did a normal amount of work in the most expensive way: the biggest model for every job, one chat that has been open for days, and a long history carried along for a question that needed none of it. It feels fine because nothing looks wrong until the "limit reached" message appears.
The rule: spend your limit on the work, not on the carrying costs. Check your usage at the halfway mark of a session, and when you are half-spent, change one of three habits: the model, the context size, or the chat.
"Half tank" is our label, not Anthropic's. It is a checkpoint, not an official threshold. The mechanics underneath it are documented.
Anthropic's help centre says usage depends on conversation length and complexity, the features you use, which model you chat with, and your effort level. Tools such as extended thinking, web search and connected apps are described as token-intensive. In Claude Code, the docs add that your full conversation is sent with every request, so a one-line question in a session that has been open all day still draws on the whole history. That is the mechanism: cost scales with context, not with how clever your question was.
Open Settings > Usage on claude.ai. Anthropic describes progress bars for your five-hour session and your weekly usage. Glance at it at natural breaks. If a bar is around half and you are not nearly done, work down this card.
| Check | Question to ask | Action |
|---|---|---|
| 1. Model | Is this job hard enough for the model I'm on? | Switch down for routine work |
| 2. Weight | Is the conversation carrying things it no longer needs? | Compress, or trim tools and files |
| 3. Chat | Has the topic changed since this chat began? | Start a fresh chat with a short handoff |
The help centre confirms models consume usage at different rates, but it does not publish a ranking, and the lineup changes. The Claude Code cost guidance gives the durable principle: the mid-tier model handles most tasks well and costs less than the top one, so reserve the top model for complex architecture or multi-step reasoning. Check the current model names in the picker rather than trusting any list written down here.
Practical sorting:
Long chats cost more per message, and Anthropic notes that where code execution is enabled, Claude summarises earlier messages automatically as you near the context limit, which itself uses more of your allowance. So compress on your terms, earlier.
/compact summarises history (you can add focus instructions, for example /compact Focus on code samples and API usage), and /clear starts fresh at no cost. Note the docs warn that compacting a large context is itself a large request, so do it before the context is huge.Start fresh when the topic changes, when the chat has been used for several unrelated jobs, or when you are returning after a long break. The Claude Code docs say the first message after a break longer than the cache lifetime reprocesses your full context (an hour on a subscription, at time of writing). A fresh chat avoids paying that on a history you do not need.
Paid plans can search past chats, and memory is on by default for Free, Pro and Max, so you often do not need to paste the old thread. When you do want continuity, use a handoff:
You are helping me close out a long working session so I can continue in a new chat with a small context.
Context: the work so far is about [TOPIC_OR_PROJECT]. I will paste the result into a new chat, so it must stand alone.
Goal: a handoff note of no more than [WORD_LIMIT] words.
Include, in this order:
1. The decision or deliverable we reached, in two sentences.
2. Facts, numbers and constraints that must not change, as bullets.
3. What is still open, as a numbered list.
4. The next single action.
5. Any instructions about tone or format I gave you.
Rules: use only what is in this conversation. If something I will need is missing or ambiguous, list it under "Questions for me" instead of guessing. Leave out discarded ideas and false starts.
Before answering, check that a stranger could continue from the note alone, and that nothing in it contradicts the conversation.
Fill in the topic and a word limit. Read the note before you paste it somewhere; it is a summary, and summaries drop things.
Priya, a freelance bookkeeper, has one chat open since Monday. It holds a client email rewrite, a spreadsheet question and a long tax-year planning discussion. On Thursday at 2pm her session bar looks about half-spent. She runs the card: the email rewrites move to a smaller model (check 1); the large pasted ledger goes into a Project instead of the chat (check 2); and she asks for a handoff note on the tax planning and opens a new chat with it (check 3). Her remaining afternoon runs in small, focused contexts.
At time of writing, Claude Pro is listed at $20 a month (about £15 to £16; check the £ price at checkout), and Max starts at $100 a month (about £75 to £80, again check at checkout) with larger usage multiples. Paid plans can also buy extra usage credits; manage these in Settings > Usage.