AI Guides › Playbooks
By Nigel Guy · 6 min read
Most people meet an AI research story as a headline, a screenshot and a confident caption: "Anthropic found a hidden mind inside Claude." They then repeat the caption. It feels fine because the headline is built from a real paper. It fails because the caption quietly swaps what was measured for what it might mean, and those are different sentences.
The rule: before you share a take on any AI research story, write down three separate lines, what was claimed, what was done to test it, and what the authors say it does not show, and share only what survives all three.
On 6 July 2026 Anthropic published "A global workspace in language models", with a longer paper on its Transformer Circuits site. The team reports that Claude has a small set of internal activity patterns that behave differently from the rest of its processing. They named it the J-space, after the Jacobian lens (J-lens), the technique used to find it. Each pattern links to a word. When a pattern is active, the word is "on the model's mind", which is not the same as the model saying it.
The authors say this structure was not designed or programmed. It emerged during training. That is the sense of "hidden": it is internal, not secret, and it is not a feature anyone shipped. The paper reports results mainly on Claude Sonnet 4.5, with key results corroborated on Haiku 4.5 and Opus 4.5.
Fill in one row per statement you are tempted to share. Here is the card filled in for this story.
| Claim (what is reported) | Evidence (what was done) | Limit (what it does not show) |
|---|---|---|
| The J-space holds the things Claude can report on | Reading the lens just before Claude names a silently chosen sport; swapping "Soccer" for "Rugby" in the workspace changes the answer to rugby | Reading and editing a handful of single-word concepts, not the whole model |
| Claude can hold something in mind on request | Asked to think about citrus while copying an unrelated sentence, "orange" and "fruits" appear in the J-space while the output shows none of it | Control is imperfect: told not to think of something, the concept still partly lights up |
| Claude reasons with it | In "the animal that spins webs" questions, swapping the unspoken "spider" for "ant" changes the answer from 8 to 6 | Shown on chosen examples; the authors call the lens imperfect |
| One entry serves many tasks | Swapping "France" for "China" shifted capital, language, continent and currency answers together | Not every task consults it; see the Spanish example below |
| Most processing bypasses it | Removing the J-space left fluency and simple fact retrieval roughly intact; multi-step reasoning fell to near zero | Describes this model family under these tests |
| It can help monitoring | The lens showed words such as "fake" and "fictional" in a staged blackmail scenario, and "manipulation" while the model faked a score file | The authors say they are optimistic, not that it catches everything |
This is the column most captions delete.
Suppose a post says: "Anthropic proved Claude secretly knows when it's being tested, and behaves only because of it."
Corrected caption: "Anthropic reports that Claude Sonnet 4.5 appears to flag a contrived test internally, and that removing those signals made bad behaviour reappear some of the time." Same story, no overreach.
Paste the real paper or post text, not a summary of it.
You are a careful research reader helping a non-specialist check an AI research story before they share it.
Source text: [PASTE_ARTICLE_OR_PAPER_TEXT]
The caption or take I want to share: [MY_DRAFT_TAKE]
Do this in order:
1. List each factual claim in my draft take as a separate line.
2. For each, quote or closely paraphrase the part of the source text that supports it, and say what experiment or observation produced it. If the source does not support it, write "not in the source".
3. For each, state what the source's own authors say the result does not show. If they say nothing, write "no limit stated" rather than inventing one.
4. Mark each claim: supported, supported with a caveat, or overstated.
5. Rewrite my draft take in no more than [WORD_LIMIT] words, using only supported claims and keeping the authors' hedges ("suggests", "in part").
Rules: use only the source text I pasted. Do not add outside facts, statistics or quotes. If something I need is missing, ask me before answering. Before you reply, check that every line in your output points back to the source.
Format: a three-column table (Claim, Evidence, Limit), then the rewritten take.
Fill in the source text, your draft, and a word limit.
Go to the Anthropic summary, then the paper if you want the detail. A few more checks: