AI Guides › Step-by-step guides
By Nigel Guy · 7 min read
The usual mistake is to pick one Claude model, set it as the default for every agent you run, and then change it whenever a headline says a new one has arrived. It feels sensible because there is only one setting to remember. It fails quietly: your cheap, repetitive agents burn the dearest model, and your hardest agent gets whatever the default happens to be.
A note on the headline itself. Some posts say Anthropic has "just changed which model powers long-running agents". We could not find an official announcement worded that way. What the current docs do say is that Claude Opus 5.5 is aimed at long-running agentic coding and knowledge work, and Claude Fable 5.1 at demanding reasoning and long-horizon agentic work. This guide is built on those verifiable pages, not on anyone's private setup.
The rule: give each agent a role, give each role a model on a written card, and change a card line only when a test on your own work says so.
You need Claude Code installed and one or more recurring agents or subagents you already use (a researcher, a reviewer, a drafter, a checker). Everything below uses documented Claude Code settings.
What it costs, at time of writing (check the live pricing page, and your £ price at checkout or in your cloud bill):
| Model | API ID | API price per million tokens (input / output) | Roughly, in £ |
|---|---|---|---|
| Claude Fable 5.1 | claude-fable-5-1 |
$10 / $50 | about £7.50 / £37 at time of writing |
| Claude Opus 5.5 | claude-opus-5-5 |
$4 / $20 | about £3 / £15 at time of writing |
| Claude Sonnet 5.5 | claude-sonnet-5-5 |
$2 / $10 | about £1.50 / £7.50 at time of writing |
| Claude Haiku 4.5 | claude-haiku-4-5 |
$1 / $5 | about £0.75 / £3.75 at time of writing |
The £ figures are rough conversions and move with the exchange rate. If you use a Claude subscription instead of the API, these per-token prices do not apply to you. Your plan's included limits do, and Fable use can bill to usage credits instead of your plan limits depending on your plan and seat tier.
On paper, list each agent and answer three questions: what is its job, what does a bad answer cost you, and how long does it run unattended? Then fill in a card like this. The first column is yours; the model column follows the docs' own descriptions.
| Role | Bad answer costs | Model to try first | Why |
|---|---|---|---|
| Long unattended build or refactor | High | Opus 5.5 | Docs position it for long-running agentic coding |
| The one hard problem Opus 5.5 keeps missing | High | Fable 5.1 | Docs suggest it when your evals on Opus 5.5 at higher effort still fall short |
| Everyday coding and drafting | Medium | Sonnet 5.5 | Docs: best combination of speed and intelligence |
| Quick lookups, tagging, summarising | Low | Haiku 4.5 | Docs: fastest model |
Claude Code's own alias descriptions agree in spirit: fable for the hardest, longest-running tasks, opus for complex reasoning, sonnet for daily coding, haiku for fast, simple tasks.
In Claude Code, type /status. It shows the active model and account details. Do this first, because the default is not what many people assume. At time of writing, the documented default is Opus 5.5 on Pro, Max, Team, Enterprise and the Anthropic API, and Fable is not the default on any plan; you must choose it.
Pick one:
/model for the picker. Press Enter to switch and save it as your default, or s to switch for this session only. Typing /model sonnet directly behaves like Enter, so it saves the default.claude --model opus (or sonnet, haiku, fable). This affects only that session."model": "opus" in ~/.claude/settings.json.Priority runs from the --model flag, to the ANTHROPIC_MODEL environment variable, to settings. Resumed sessions keep their original model unless it was retired.
For a plan-then-build split in one session, the opusplan alias uses Opus for planning and Sonnet for execution.
Open the subagent file and add a model line to its frontmatter:
---
name: code-reviewer
description: Reviews code for quality and best practices
tools: Read, Glob, Grep
model: sonnet
---
Allowed values are the aliases sonnet, opus, haiku or fable, a full model ID such as claude-opus-5-5, or inherit to follow the main conversation. If you leave it out, Claude Code resolves the model in this order: the per-invocation model parameter, then the frontmatter, then the CLAUDE_CODE_SUBAGENT_MODEL environment variable, then the main conversation's model. To force every subagent onto one model, set CLAUDE_CODE_SUBAGENT_MODEL together with CLAUDE_CODE_SUBAGENT_MODEL_FORCE to 1; that overrides each agent's own line, so use it for cost tests, not as a permanent setting.
Models differ in default effort. The API docs list Opus 5.5 at medium and Sonnet 5.5 and Fable 5.1 at high. Set it explicitly with /effort high, the CLAUDE_CODE_EFFORT_LEVEL variable, or effortLevel in settings (there is also a per-model modelSettings block). Before blaming a model, check you are comparing like with like.
Aliases move as new models arrive. To stop an agent changing underneath you, pin them in settings or your shell:
{
"env": {
"ANTHROPIC_DEFAULT_OPUS_MODEL": "claude-opus-5-5",
"ANTHROPIC_DEFAULT_SONNET_MODEL": "claude-sonnet-5-5",
"ANTHROPIC_DEFAULT_HAIKU_MODEL": "claude-haiku-4-5"
}
}
Note that the docs list Haiku 4.5 retirement as not sooner than 15 October 2026, so check the deprecations page before you pin it for a long project.
Use this prompt to build the test, then run it yourself.
You are an evaluation designer helping me choose a Claude model for one agent role.
Role: [AGENT_ROLE]
What the agent does in one sentence: [JOB_DESCRIPTION]
Models on my card for this role: [CURRENT_MODEL] and [CANDIDATE_MODEL]
What a good result looks like: [SUCCESS_CRITERIA]
What a bad result costs me: [COST_OF_FAILURE]
Three real past tasks I can reuse: [TASK_1], [TASK_2], [TASK_3]
Do this in order:
1. Ask me for any missing input above. Do not guess.
2. Propose a pass/fail checklist of no more than six items, written before any run.
3. Give me a table with one row per task and columns for each model: passed checklist, retries needed, time taken, my notes.
4. Tell me how to run each task identically on both models, including effort level.
5. After I paste results, recommend keep, switch or retest, and say what evidence would change your mind.
Self-check before replying: is every checklist item observable without opinion, and did you avoid inventing any benchmark figures?
Fill in the square-bracket fields with your own role, models and three past tasks.
/status shows the model you intended, in a fresh session.model: haiku and see it behave accordingly, or check your usage records)._FORCE variable once the test is over.-p or the Agent SDK) do not ask for consent before Fable usage bills to usage credits. Set a spending limit before you automate.