AI Guides › Playbooks

The Condition Card: Reading GPT-6 Astra's Benchmark Numbers

By Nigel Guy · 6 min read

The usual way to read a launch chart is to find the biggest bar, remember the percentage and carry it into a purchasing decision. That feels efficient, and it is how a 99.9% on one test turns into "it solves everything" by the time it reaches a group chat. The numbers on OpenAI's GPT-6 Astra page are real, in the sense that OpenAI published them. What they measure depends on the conditions attached, and the conditions sit in footnotes and captions.

The rule: never carry a benchmark number anywhere without its conditions. If you cannot say what harness, effort setting, safeguards and rival scores sat behind it, you are repeating a headline, not a result.

The mechanism: the Condition Card

For any score you plan to rely on, fill in five lines. If a line is blank, the number is not ready to use.

Line Question Where to find it
1. Who ran it The vendor, an independent lab, or you? Footnotes, system card, third-party leaderboards
2. Harness What scaffolding, tools and settings surrounded the model? Footnotes
3. Effort Which reasoning setting produced the score? Chart notes
4. Safeguards Is this the model you can actually use, or a less restricted one? Safety sections
5. Comparison Who is on the chart, and who is missing or modified? Table dashes and footnotes

On OpenAI's page, the evaluation note says scores are "the maximum at any effort", and that GPT evaluations were run in OpenAI's research environment or API, which "may provide slightly different output" from production ChatGPT. Line 1 and line 3 are answered for every number on the page before you start: OpenAI ran them, at the best setting it tried.

Four numbers, four cards

All figures below come from OpenAI's launch page as published at time of writing. They are the vendor's own results unless stated.

1. ARC-AGI-3: 99.9%

The headline. The comparison column shows GPT-5.6 Sol at 7.8% and Claude Opus 5 at 30.2%. Footnote 1 says Astra was run with OpenAI's Responses API harness, which "changes two settings to better match real-world performance", and that the changes do not specifically target ARC-AGI-3. OpenAI also quotes the ARC Prize Foundation saying Astra beat its human action-efficiency baseline on 96% of levels.

Card verdict: strong, but the number belongs to a harness configuration, not to the bare model. Do not read it as what you get from a casual chat prompt.

2. ExploitBench: 100%

OpenAI states this was measured "without production safeguards". The launch version, it says, will refuse more advanced cybersecurity tasks such as writing proof-of-concept exploits, with wider access planned through its Daybreak programme. OpenAI also says some tasks in its newer June-August benchmark may not permit full success, and reports 39.0% there.

Card verdict: this is a measurement of capability, not of what you can do with the model today. Line 4 is the whole story.

3. Terminal-Bench 4.0: 57.9%

Astra scores 57.9% against 55.8% for Claude Fable 5.1, 52.6% for Claude Opus 5 and 37.3% for GPT-5.6 Sol. OpenAI adds that its estimated API cost per task is about 9% lower than Sol and 63% lower than Fable 5.1. The gap to Fable 5.1 is about two points. Treat that as "competitive" rather than "decisive", and note that the cost claim is also an OpenAI estimate.

Card verdict: useful for coding-agent work. The cost line may matter more than the score line.

4. The Artificial Analysis Intelligence Index: 61.2

The least-quoted row. OpenAI's own table lists Astra at 61.2, Fable 5.1 at 65.7, Opus 5 at 63.1 and Fable 5 at 62.1. On that composite index Astra is not first, and OpenAI printed it anyway.

Card verdict: a reminder that "best on the benchmarks I chose to feature" and "best overall" are different claims.

What it is actually good for

Going by OpenAI's own evidence, the strongest cases are computer use (Agents' Last Exam 59.3%, OSWorld 2.0 72.6% with a footnoted offline subset), terminal and coding agents, long-context recall (MRCR 8-needle 100% at 256K-512K) and heavy maths. Your own task decides whether any of that transfers.

Some comparisons come with caveats you should carry. Claude scores on BenchCAD reflect modifications listed in Anthropic's system card, and Claude Fable models are omitted from several science benchmarks because, OpenAI says, they refuse most questions. A dash in a table is information, not a tie.

Access and price, at time of writing: OpenAI says Astra rolls out to Plus, Pro, Business and Enterprise users, with Enterprise access off by default, and the API model is gpt-6-astra. API pricing is $10 per million input tokens and $50 per million output tokens (about £7.50 and £37 at time of writing; check your own £ figure at billing). Fast mode is up to 2x the speed at 2x the price. OpenAI's models page also lists GPT-6.1 Sol as "near-Astra performance" at a lower cost, listed at $2 and $10 per million tokens (about £1.50 and £7.50 at time of writing). Check the pricing page before budgeting; these move.

A worked example (hypothetical)

Imagine a three-person UK agency choosing a model for an agent that updates client records in a CRM. Someone shares the OSWorld 2.0 number, 72.6%.

Verdict: a 2.4-point gap on a partial, vendor-run score is not a reason to switch. The agency runs ten of its own CRM tasks on two models, scores them pass or fail, and records cost per completed task.

A prompt to run the card for you

You are a sceptical analyst who checks AI benchmark claims. I will give you a claim and the source text around it.

Claim: [BENCHMARK_CLAIM]
Source text (footnotes, captions, tables): [PASTED_SOURCE_TEXT]
My use case: [MY_USE_CASE]

Goal: fill in a Condition Card with five lines: who ran it, harness, effort setting, safeguards, comparison set. Quote the exact words from the source for each line. If the source does not say, write "not stated" and do not guess.

Then give me:
1. A one-sentence plain-English reading of what the number does and does not show.
2. Whether it plausibly transfers to my use case, with the reason.
3. One cheap test I could run myself, with a pass threshold set in advance.

Before answering, check that every statement is backed by the pasted text. If I left a placeholder empty, ask me for it before starting.

Fill in the three bracketed items; paste the page text rather than a link, so it works from what you provide.

What to skip

Guardrails

Sources

All 751 AI guides · JulieMango plans from £17/mo