AI Guides › Playbooks
By Nigel Guy · 6 min read
Most people hear "o3 is being retired", pick whatever the newest model is, paste in their old prompts and carry on. It feels fine because the new model answers politely. The failure shows up weeks later, when a saved prompt that used to follow your format now ignores it, or a script that quietly depended on a pinned snapshot stops returning anything at all.
The rule: find out which o3 you actually use, put a date against it, then swap on a test card, not on faith.
"o3" means two different things depending on where you use it, and the dates differ. This is where most of the confusion comes from.
| Where you use it | What OpenAI says | Date |
|---|---|---|
| ChatGPT (model picker) | o3 retired from ChatGPT after a 90-day sunset period | 26 August 2026 (already passed) |
API, snapshot o3-2025-04-16 (the alias o3 points here) |
Deprecated, replacement listed as gpt-5.6-sol |
Shutdown 11 December 2026 |
API, o3-pro-2025-06-10 |
Deprecated, replacement listed as gpt-5.6-sol with pro mode |
Shutdown 11 December 2026 |
API, o3-deep-research, o4-mini, o4-mini-deep-research |
Deprecated 22 April 2026 | Shutdown 23 July 2026 (already passed) |
Two things to hold onto. First, the ChatGPT retirement is the one that has already happened; the API date is the one still ahead of you. Second, if a tool you use calls the API on your behalf, it may be affected even though you never typed "o3" anywhere.
The ChatGPT date comes from OpenAI's announcement as reported by secondary sources; I could not open OpenAI's own help-centre release notes to read it first-hand, so check the model release notes in the help centre if the exact date matters to you. The API dates and replacements are from OpenAI's deprecations page, which I read directly.
One card per saved prompt or workflow. Four boxes, filled in this order.
List every place o3 could be hiding:
o3, o3-pro, o4-mini and any dated name such as o3-2025-04-16.Write down, for each: the exact model string, who owns it, and the shutdown date from the table above.
Run three to five real inputs through the old setup now, while it still works, and save the outputs. This is your baseline. Choose one typical input, one awkward one, and one you know has gone wrong before. If o3 has already gone from ChatGPT, use any saved outputs from past chats instead.
Change one thing only: the model. OpenAI's deprecations page lists gpt-5.6-sol as the replacement for o3-2025-04-16. OpenAI's model page lists its reasoning effort settings as none, low, medium (default), high, xhigh and max. Start on the default and only raise it if the baseline comparison shows a gap, because higher effort generally means slower and more expensive answers (check the current pricing page for your own numbers; I have not quoted any, as pricing for this model was not something I could confirm from an official source).
Note that the model page does not itself say it replaces o3. The replacement claim sits on the deprecations page. Treat it as OpenAI's suggested default, not a guarantee of identical behaviour.
Rerun your Box 2 inputs and score each output on the same three questions:
| Check | Pass if |
|---|---|
| Format | The structure you asked for is intact (headings, table, JSON keys) |
| Substance | Facts, reasoning and caveats match or beat the baseline |
| Edge case | The awkward input still behaves |
For each prompt, mark one of: ship (passes), patch (fails on format or edge case, so edit the prompt), or retire (you no longer use it). Write a date by which the swap must be live. For API work, give yourself a buffer well before 11 December 2026, not on it.
When a prompt fails after the swap, ask the new model to help repair it rather than guessing.
You are a careful prompt engineer helping me move a saved prompt from one model to another.
Context: my prompt used to run on [OLD_MODEL]. It now runs on [NEW_MODEL]. It no longer behaves the same.
Here is the prompt:
[PASTE_PROMPT]
Here is an example input:
[EXAMPLE_INPUT]
Here is what the old model produced (good):
[OLD_OUTPUT]
Here is what the new model produces (not good):
[NEW_OUTPUT]
Steps:
1. List the specific differences between the two outputs (format, length, tone, missing content).
2. For each difference, name the line or missing instruction in my prompt most likely to explain it.
3. Rewrite my prompt to fix those differences. Keep my original purpose and placeholders.
4. Show the rewritten prompt in a code block, then a short list of what you changed and why.
Constraints:
- Do not invent requirements I have not stated. If something is ambiguous, ask me before rewriting.
- Do not make the prompt longer than it needs to be.
- If the outputs differ only in ways that do not matter for [MY_GOAL], say so and change nothing.
Before you answer, check that every change you made maps to a difference you listed in step 1.
Fill in the two model names, the prompt, one example input, both outputs, and a one-line goal.
Priya runs a three-person bookkeeping firm. She has a saved ChatGPT habit of picking o3 to summarise client email threads into a table, and one script her freelancer wrote that calls o3-2025-04-16. Her card: the ChatGPT habit is already gone (26 August), so it is a "patch" job using old saved outputs as the baseline. The script is a "ship if passes" job with a 11 December deadline, so she sets her own date in early November. She tests five real threads; four pass, and the awkward one, a thread with forwarded attachments, loses a column. She runs the patch prompt above, adds one line about forwarded content, retests, and marks it ship.