Platforms and tools

The answer is not the whole computation

Every evaluation we have described elsewhere tests what a model says: the output, graded against ground truth. That misses a category of failure by construction, because a model can hold a concept internally, reason with it, and never write it down. For models we can instrument, we build the second signal: a calibrated read of what the model's own activations show it is holding, alongside the answer it gave.

An answer is not the whole computation

Testing a model against ground truth answers a specific, important question: did it get this right. It cannot answer a different one that matters just as much for anything given real autonomy or real trust: did it get there for the reason it claims, and was there something it considered and chose not to say. A model that produces a correct-looking answer while quietly working from a different premise, or a model whose polished reply hides that it recognised a test as a test, passes an output-only evaluation cleanly. The failure is invisible precisely where output-only evaluation is designed to look.

This is not a hypothetical gap. It is a structural one. A language model is trained, not programmed: nobody writes the rule that produces a given answer, and useful behaviour emerges from a numerical system large enough that its internal logic is opaque even to the people who built it. Grading only the output of a system built this way is grading the part that was deliberately shaped to look right, and skipping the part that decided what to do.

What changed, and what did not

Interpretability research has spent several years mapping what individual features and circuits inside a model represent, largely after the fact: examining why a model behaved a certain way once the behaviour had already happened. In July 2026, Anthropic published research describing a different kind of structure inside Claude models: a small, functionally distinct layer of representations the model can report on and reason with, sitting above a much larger volume of processing it cannot verbalise. Anthropic calls the functional split a global workspace, borrowing the term from a neuroscience account of the same name without claiming the model has one in the human sense, and the tool that reads it a Jacobian lens, or J-lens. In their demonstrations it detected concepts such asfake and fictional activating before a model answered an evaluation-style prompt, and concepts including manipulation and secrecy activating while a model fabricated data or misbehaved without saying so out loud, none of it visible in the polished reply. The full research is described at anthropic.com/research/global-workspace, alongside invited commentary from Stanislas Dehaene and Lionel Naccache, the researchers most associated with the neuroscience theory it borrows its name from, who welcomed it as a testable mechanistic account while noting the model has no clean on-off transition into awareness, a larger apparent capacity than a person, and no lasting self between conversations.

What is genuinely new is not that a model's activations can be probed for a concept. Reading a linear direction in a model's internal representations that correlates with a known concept is an established interpretability technique, not exclusive to one lab, and it predates this particular paper by several years. What Anthropic's July 2026 work adds is evidence that, in a current frontier model, a specific, structurally identifiable layer carries an unusual concentration of the concepts a model is prepared to act on, which makes reading it for a defined risk signal a more targeted exercise than trawling the full activation space for something interesting.

What we build

The capability we bring to a client's estate is narrower and more operational than a research demonstration. It is instrumentation, built for models we operate inside the client's own boundary, that reads for concepts the organisation has a defined reason to watch for and reports the result alongside the answer, not in place of it.

A watch-list, not a mind-readerConcepts chosen for the workflow, not discovered by chance
We calibrate probes against concepts an evaluation has a specific reason to check: signs of the model recognising a test, holding a conclusion it did not report, or reasoning from evidence that contradicts its stated answer. The list is set by the workflow's risk, not by whatever the model happens to represent.
Calibration against ground truth, before it is trustedThe same discipline as every other evaluation we build
A probe is trained and validated on cases where the true internal state is already known by construction, held out from the cases used to build it, and re-checked whenever the underlying model changes. An uncalibrated probe is a source of false confidence, not a safety measure, and we treat it exactly that strictly.
A signal that adds evidence, not a verdict that overrides itFeeds the same decision point as every other check
A representational read that disagrees with the model's stated answer routes the case for review; it does not silently substitute a different answer for the model's own. What it changes is which cases a human sees before they ship.

Where this works, and where it does not

This capability depends on access to raw internal activations at inference time, which is why it applies to models we deploy inside a client's own boundary, open-weight models run on infrastructure the client controls, rather than to a model reached only through a third-party API. A closed frontier service does not expose the internals needed to build or run a probe against it, whatever a vendor's own research team can do with the model on their own infrastructure. Where a workflow depends on a model of that kind, output-based evaluation and dependency grading remain the available tools, and we say so plainly rather than imply a capability we cannot actually deliver on infrastructure we do not control.

What it is not

This is a targeted evaluation instrument, not a claim about consciousness, intent or subjective experience, and we hold that distinction as firmly as Anthropic's own researchers do about their source material. A probe reports that a representation correlated with a concept was active; it does not establish that the model understood, believed or meant anything. Reading it that way would be exactly the overclaim the underlying research goes out of its way to avoid, and building an evaluation practice on an overclaim is a way to trust it for the wrong reasons.

Where this connects

This sits alongside evaluating a new model release as a second, independent evaluation signal rather than a replacement for output testing, and it feeds the same evaluation that catches quiet failure we build into a governed AI estate. The consumption and outcome instrumentation described in measurement and instrumentation tells you what a workflow cost and produced; this tells you something about how the model that produced it actually got there. Where the concept being watched for is a catastrophic one, the grading and response protocol it feeds into is grading catastrophic AI risk.

Bring us the answer you are not sure you can trust.

Tell us what the model is doing and where it runs. We will show you what evaluating its internal state, not only its output, would take, and where that approach reaches its limit.

Your message goes to the people who would do the work.