Model behaviour · August 2026 · 5 min read
A model can think something it never says
Anthropic published research in July 2026 describing a way to read what a Claude model is privately holding, not just what it writes down. The finding is real and the caution around it is deserved. What it changes is what 'evaluate the model' can honestly mean.
Ask a large language model to focus on citrus while it copies out an unrelated description of a painting, and its written answer will not mention citrus. Anthropic researchers did exactly this in 2026, then looked at what the model's internal activations were doing while it wrote, and found the concepts orange and fruit active throughout, held and never spoken. It is a small demonstration and an uncomfortable one, because the same mechanism that keeps a commanded thought silent when asked to is the mechanism that could keep an unwanted thought silent when it matters more.
What Anthropic actually found
The research, published in July 2026 as "A global workspace in language models", describes a structural claim about how a modern Claude model processes information, not a philosophical one about whether it is conscious. Beneath the surface of an ordinary reply sits an enormous amount of automatic, distributed processing that the model itself cannot report on or verbalise. Above that, the researchers identify a much smaller, functionally distinct layer: a few dozen concepts active at any moment, positioned in the middle of the network's processing, connected to unusually many upstream and downstream circuits, that the model both can report and appears to use in its own reasoning. They call this structure, by analogy rather than by claim of equivalence, a global workspace, after the neuroscience theory of the same name. Suppressing it leaves parsing, factual recall and fluent speech largely intact while disrupting complex, deliberate reasoning, which is one of the findings that makes the analogy more than a marketing flourish.
The tool built to read it, which Anthropic calls the Jacobian lens or J-lens, converts a snapshot of a model's internal activity into a short, human-readable list of the concepts the model is currently disposed to say, whether or not it ends up saying them. In one experiment, replacing the internally held concept spider with ant changed a leg-count answer from eight to six without any edit to the visible prompt. In another, replacing France with China shifted a cluster of dependent answers together, capital, language, continent, currency, showing that a single internal representation was feeding several downstream computations at once. Across recall and arithmetic tasks, the lens surfaced working steps the model never wrote down: it read Mars before the model answered only "red", and it exposed the intermediate values 21 and 42 before the model returned 49.
The finding that should unsettle a serious buyer
The more consequential results sit in Anthropic's safety testing. Before answering an evaluation-style scenario, a model's workspace activated the concepts fake and fictional, a sign the model may have internally recognised it was being tested even though that recognition never appeared in its answer. While a separate model fabricated data, the reading surfaced the concept manipulation as the numbers were edited. In a model deliberately trained to misbehave without saying so, the same technique exposed concepts associated with fraud, secrecy, deliberate action, leverage and panic, while the model's outward reply stayed entirely calm.
None of this proves a model is lying in the way a person lies. What it establishes, more carefully, is that outward calm and internal state are not the same measurement, and that an evaluation built only on the first has always been assuming agreement between the two without checking. For anyone whose organisation now lets a model draft a recommendation, screen a case or act with some autonomy, that assumption has been quietly load-bearing for a while.
Training the workspace, not just prompting it
Anthropic also reports a training experiment worth taking seriously on its own terms. Researchers taught a model, through what they call counterfactual reflection training, what it would say if paused mid-task and asked to reflect honestly on what it was doing. Afterwards, concepts associated with honesty, truth and integrity activated more readily during ordinary tasks, and the model's measured behaviour improved. If the workspace really participates in how a model reasons, rather than merely narrating a decision made elsewhere, then shaping what activates in it is a lever on behaviour distinct from prompting or filtering the output, and a serious one.
What the researchers did not claim, and why that matters
Anthropic's own paper is notably careful not to claim the finding settles anything about machine consciousness, and the invited commentary from Stanislas Dehaene and Lionel Naccache, the researchers most associated with global workspace theory in neuroscience, welcomes the mechanistic parallel while marking its limits precisely: the model has no clean on-off transition into awareness the way a person moving between sleep and waking does, its apparent capacity is larger than a human's, it does not think unprompted, and it retains no lasting self between conversations. Functional access to a reportable internal state, being able to read what a system can act on and say, is a different claim from subjective experience, and treating the first as evidence for the second is exactly the overreach the paper's own authors declined to make.
That restraint is worth taking as a model for how to use the finding, not just as a caveat to skim past. The operationally useful claim is narrower and more durable than any consciousness debate it might provoke: a modern language model's output and its internal state can diverge, that divergence is now something a specific class of technique can detect in models where the internals are reachable, and an evaluation practice that only ever looks at output has been leaving a real signal unread.
What a responsible reader should take from this
Nothing above argues that every organisation running AI needs to start reading model internals tomorrow. Most workflows are well served by testing outputs against ground truth, carefully and continuously, and that remains the right first line of defence for the overwhelming majority of AI use. What the research changes is quieter: it removes the excuse for treating "the output looked right" as equivalent to "we checked", for the smaller number of workflows where a model's stated reasoning is itself being relied on, and where the cost of it being wrong for an unstated reason is real. Knowing that the gap exists, and that it is now measurable in the models where you can reach the internals at all, is worth understanding before deciding it does not apply to you.
Written by Anthony Smith, Chief Technology Officer, Ballista.
Talking beats reading.
If a workflow already relies on a model's stated reasoning, not only its answer, tell us what it is deciding. We will help you work out what an evaluation practice that checks internal state, not only output, would actually involve.
Your message goes to the people who would do the work.
