AI Guides › Workbench
By Nigel Guy · 7 min read
Most people judge a new model by reading the launch table, picking the biggest name and chatting with it until it feels clever. That tells you what the vendor chose to measure, not whether the model earns its price on your jobs. GPT-5.6 is a good case, because it ships as three tiers and the cheapest may do your work for a fraction of the cost.
The rule: run one real job of yours through all three tiers against a pass mark you wrote down first, and pay for the cheapest tier that passes.
A dating note before anything else. GPT-5.6 reached general availability on 9 July 2026, so it is no longer "just live". OpenAI's own site now lists a GPT-6 family, with GPT-6.1 Sol announced on 29 September 2026. Nothing below depends on GPT-5.6 being the newest; the Test Card works on any tier ladder. Check which models your plan currently offers before you start.
OpenAI says the number is the generation and Sol, Terra and Luna are durable capability tiers that can advance on their own cadence. Figures below come from OpenAI's launch post and its 30 July price update. Prices are per million tokens and the £ figures are rough conversions, about £X at time of writing.
| Tier | What it is | API price, launch | API price after 30 July | Best for | Catch |
|---|---|---|---|---|---|
| Sol | Flagship; adds max effort and ultra (four parallel agents by default) |
$5 in / $30 out (about £4 / £22) | Unchanged on 30 July; OpenAI later announced a cut of over 20% for three months from 21 August, so check the current page | Hard, long-running professional, coding or research work | Most expensive; ultra trades higher token use for speed |
| Terra | Balanced everyday model, described as competitive with GPT-5.5 | $2.50 in / $15 out (about £1.85 / £11) | $2 in / $12 out (about £1.50 / £9) | Routine drafting, analysis, workspace Q&A | Below Sol on the hardest tasks in OpenAI's own table |
| Luna | Fastest, cheapest | $1 in / $6 out (about £0.75 / £4.50) | $0.20 in / $1.20 out (about £0.15 / £0.90) | High-volume, well-specified work, background automation | Weakest of the three on long-context and abstract reasoning in OpenAI's table |
In ChatGPT, access varies by plan. At launch, Free and Go users got Terra in ChatGPT Work and Codex; Plus, Pro, Business and Enterprise could choose all three. An August update made Luna the default for Free and Go in Chat, with a Think button for harder questions. Plans change often, so confirm in the model picker.
Read this with care: every number is OpenAI's own, from its launch table, and the Claude models are named as that table names them (Fable 5, Mythos 5, Opus 4.8). I have not checked them against Anthropic's own reports.
The honest summary is "competitive, with different strengths", which is exactly why you test on your own work.
Fill this in before you open any model.
| Field | What you write |
|---|---|
| Job | One real task you repeat, with its real inputs |
| Pass mark | 3 to 5 checkable criteria, written now |
| Cost cap | The most you will pay per run |
| Runs | Same brief, same inputs, all three tiers |
| Verdict | Cheapest tier that meets every criterion |
You are helping me design a fair test of three AI model tiers.
Context: I want to test [JOB_DESCRIPTION] on a cheap, a mid-range and a flagship model.
My inputs look like: [INPUT_DESCRIPTION].
Who uses the output and what they do with it: [AUDIENCE_AND_USE].
Task: Write a Test Card with (1) one precise job statement, (2) four or five pass-mark
criteria that I can check by eye without trusting the model, (3) one trap that a lazy
answer would fall into, and (4) what evidence I should keep from each run.
Rules: Make every criterion yes/no. If you lack information, ask me up to five questions
before drafting. Do not invent facts about my business.
Before answering, check that no criterion depends on opinion.
Fill in the three bracketed fields.
OpenAI positions the tiers around agentic work, so brief them like a colleague with a deadline.
Role: You are a careful [ROLE, e.g. analyst or editor] working for [ORGANISATION].
Job: [JOB_STATEMENT].
Inputs provided: [PASTED_OR_ATTACHED_MATERIAL].
A good result: [PASS_MARK_CRITERIA].
Steps: 1) List what you understand the job to be and any missing inputs, and stop to ask
if something essential is missing. 2) Do the work. 3) Check your result against each
criterion and mark it met or not met.
Output format: [FORMAT, e.g. a one-page memo with a table of sources].
Constraints: Use only the material supplied; label anything you infer; do not guess
numbers or quotes.
Fill in the role, organisation, job, inputs, criteria and format. Run it unchanged on Luna, Terra and Sol. If you use effort settings, keep them the same across tiers first, then raise Sol alone as a separate trial.
You are a strict reviewer. Below are the pass-mark criteria and three outputs labelled A, B and C.
Criteria: [PASS_MARK_CRITERIA]
Outputs: [OUTPUT_A] [OUTPUT_B] [OUTPUT_C]
For each output, mark every criterion met or not met with a one-line reason quoting the
text. Do not guess which model wrote which. Finish with a table and name the outputs that
meet all criteria. If you cannot judge a criterion from the text, say so.
Strip the model names before pasting. Better still, mark it yourself and use this only as a second opinion, because a model grading a model is itself a bias.
A hypothetical example. You summarise a 40-page supplier contract for a small business, with pass marks: every payment date captured, every termination clause listed, no invented clauses, under 400 words, uncertainty flagged. Luna passes four of five and misses a clause in an appendix; Terra and Sol pass all five. Terra is your answer. If only Sol had caught the appendix clause, you would pay for Sol on contracts and keep Luna for first-pass triage. This is a made-up scenario, not a result.
ultra until a single Sol run has failed your card. It spends more tokens by design.I could not independently confirm Anthropic's own scores for the Claude models quoted, or the exact current Sol price after the August discount.