AI Guides › Workbench
By Nigel Guy · 7 min read
Most people treat Claude's thinking settings like a "genius mode" switch: flick it on for everything, watch the reasoning panel scroll, and assume the answer must be better because it took longer. That's confidence theatre. On current Claude models thinking is adaptive, so the model already decides per message whether to reason and how hard. On some models you can't switch it off at all. What you actually control is effort, and on the wrong task, more effort just burns your usage allowance and slows you down.
The rule: leave effort at the default for routine work, raise it only for tasks where a wrong answer is expensive and the reasoning can be checked, and judge the result by checking it, not by how long Claude thought.
| Control | What it does | Cost / availability (at time of writing) | Best for | Catch |
|---|---|---|---|---|
| Effort in the Claude apps | Sets how readily and how deeply Claude reasons: Low, Medium, High, Extra high, Max | Available on all plans. Higher levels use up your plan's usage limits faster | Raising depth for one hard conversation | It applies from the next reply on, and it stays set until you change it back |
| Thinking toggle in the Claude apps | Shows or hides the expandable reasoning section | All plans, on supported models | Older models where thinking can still be turned off | You can't turn it off on Sonnet 5.5, Opus 5.5, Fable 5.1 or Opus 5 |
/effort in Claude Code |
Sets effort per model, for the session or as your saved default | Comes with the Claude Code access on your plan | Long debugging, refactors, architecture decisions | The default depends on the model: Medium on Opus 5.5 and Sonnet 5.5, Extra high on Opus 4.7 |
ultrathink keyword in Claude Code |
Asks for deeper reasoning on one turn only | No setting change | A single hard step in an otherwise routine session | It's an instruction inside the prompt. It doesn't change the effort level sent to the API |
API output_config.effort with adaptive thinking |
Sets effort per request. Claude decides per request whether to think | Thinking tokens are billed as output tokens | Developers building products | Changing effort mid-conversation invalidates the prompt cache |
| Per-message wording | Phrases like "Please think hard before responding" nudge one turn | Free | Mixed conversations | It depends on the exact wording, so test it before you rely on it |
On price: Anthropic's pricing page lists Pro at $20 a month (about £15 at time of writing) billed monthly (or $17 a month, about £13, on an annual plan) and Max from $100 a month (about £75), in US dollars. Your checkout shows the sterling price, including VAT. Check the page before you upgrade just to get "more thinking". You don't need to, because effort and thinking are available on every plan.
In the Claude apps, click the model name next to the send button and hover over Effort. Pick a level there, and use the Thinking (or Extended) toggle where your model offers one. Anthropic's help centre suggests High as the balanced default. Low and Medium suit routine tasks and save usage. Extra high is meant for complex coding, and Max for work where correctness matters most.
Turn it up when all three of these are true:
If any of the three is missing, the default is fine.
In Claude Code, type /effort high (or /effort for the slider). Press s in the picker to apply it to the current session only, or Enter to save it as your default. /effort auto puts the model back on its default. For a single hard turn, put ultrathink anywhere in your prompt instead of changing the session. Claude Code only recognises that keyword. Plain "think hard" goes through as ordinary text.
| Raise effort | Leave at the default |
|---|---|
| Debugging something that has already beaten one attempt | Rewording an email |
| Financial or statistical reasoning with several steps | Summarising a document you'll read anyway |
| Planning a change across many files | Brainstorming names or headlines |
| Checking an argument for gaps | Looking up a fact (thinking doesn't add knowledge) |
| Comparing options against several constraints at once | Formatting, translation, tidying tables |
Use this prompt when you've decided a task needs the depth. Fill in the bracketed parts:
You are a careful analyst working on a problem where a wrong answer is costly.
Task: [DESCRIBE THE PROBLEM IN ONE OR TWO SENTENCES]
Context and materials: [PASTE DATA, CODE, CLAUSES OR NOTES]
Constraints that must hold: [LIST THEM]
What a good answer looks like: [E.g. "a recommendation with the working shown"]
Before answering:
1. If any input above is missing or ambiguous, list your questions and stop. Do not guess.
2. Work through the problem step by step, and say each assumption you make.
3. Try at least one alternative approach or reading, and say why you rejected it.
Output format:
- Answer (three sentences at most)
- Working (numbered steps)
- Assumptions (bulleted, each marked "stated by user" or "my assumption")
- What would change this answer (two or three bullets)
Self-check before sending: recheck every number and every constraint against the materials, and remove any claim you can't trace back to them.
How long it took or how long the reasoning panel is doesn't tell you much. Here's what does:
Ctrl+O (Option+O on macOS) for verbose mode, which shows the reasoning in grey italics.usage.output_tokens_details.thinking_tokens in the response. That figure counts the raw reasoning tokens you were billed for, and it's often higher than the summarised text you see. With adaptive thinking, a turn may have no thinking block at all. That's normal.Is it measurable or just a feeling? The token count can be measured. Whether the answer is better can only be measured by checking it. Anthropic's own advice to developers is to run a sample of real tasks with and without a change and compare quality, tokens and latency. Do a small version of that yourself: give the same three tasks to Medium and to High, then check which answers were actually correct.
This prompt asks Claude to grade its own reasoning. Paste the answer and its reasoning in:
You are reviewing a reasoning trace for quality, not style.
Original question: [PASTE QUESTION]
The answer and visible reasoning: [PASTE BOTH]
Check:
1. Did the reasoning consider at least one alternative? Quote it, or say "none".
2. List each assumption, and whether it was stated or hidden.
3. Find the weakest step and explain why it's weak.
4. Is there a claim with no support in the materials? List it.
If the reasoning wasn't supplied, say so and stop rather than inferring it.
Finish with one line: "Deeper effort helped / did not help / can't tell", and give the reason.
ultrathink or a session-only /effort change for hard steps.