Field note · AI & Technology

The Interpretability Breakthrough That Quantifies Its Own Limits

Anthropic's J-space finding gives auditors a window into model reasoning. The paper's own data shows that window covers less than 10% of computation.

AC
Avi Chudasama
Co-Founder & CEO, Newtral
Published
August 6, 2026
Last reviewed
August 6, 2026
Read time
7 min · 1,388 words
Current

By Avi Chudasama

On July 6, 2026, Anthropic's interpretability team published a paper - "Verbalizable Representations Form a Global Workspace in Language Models" - that introduced a technique for reading what a language model computes beyond its visible output. The technique, called the Jacobian lens, identifies a small privileged subspace of internal representations the researchers call the J-space. It is a genuine advance: for the first time, auditors have a mathematically grounded tool that can surface a model's hidden reasoning before it reaches the output. But the paper's own measurements reveal a constraint most coverage has not addressed. The J-space holds approximately 25 active concepts and accounts for less than 10% of the model's total activation variance. The breakthrough that makes AI auditing possible simultaneously quantifies how partial that audit is.

01

What the J-Lens Finds

The Jacobian lens works by computing, for every word in the model's vocabulary, the internal activity pattern that makes the model more likely to produce that word at a future position. The collection of these patterns forms the J-space - a zone of internal activity where the model holds concepts it can report on, reason with, and direct at will. The J-space was not designed or programmed by Anthropic's engineers. It emerged during the model's training process, organising itself without architectural intervention.

The paper demonstrates that the J-space satisfies five functional properties that neuroscientists have long associated with conscious access in humans: verbal report (the model names concepts represented in the J-space when asked what it is thinking about), directed modulation (swapping one J-space vector for another redirects downstream reasoning - replacing "France" with "China" causes every downstream circuit to return China's corresponding answers), unverbalized internal reasoning (the workspace carries reasoning steps the model does not write out), flexible cross-task generalization (the same workspace representations transfer across unrelated tasks), and selectivity (the workspace mediates deliberate reasoning but not automatic pattern-matching).

The parallel is to Global Workspace Theory, first proposed by cognitive scientist Bernard Baars, which holds that consciousness involves a processing hub that integrates and broadcasts information for use in reasoning, behaviour control, and speech. Anthropic's researchers present evidence that "an analogous functional distinction has emerged in modern AI models" - a workspace that broadcasts to downstream processes, surrounded by a much larger volume of automatic processing it cannot access or articulate.

02

The Ablation Test

The paper's most operationally significant experiment is an ablation study. When the J-space was suppressed entirely, tasks requiring flexible reasoning - multi-hop inference, analogy completion, translation, sonnet writing - collapsed in performance to below the level of Anthropic's much smaller Haiku model. Tasks involving factual recall, sentiment analysis, and grammatical judgment survived largely intact.

This is a clean separation. The workspace does not contribute equally to everything the model does. It mediates precisely the activities that require composition, inference, and flexible problem-solving, while leaving automatic processing - classification, recall, pattern-matching - to the 90% of the model that operates outside it.

A specific finding sharpens the picture: math problems solved with explicit chain-of-thought reasoning proved far more robust to J-space ablation than the same problems answered directly. The researchers interpret this as the model externalising onto the page what it would otherwise carry in the J-space - a strategy that mirrors how humans use scratch paper to offload working memory.

03

External Validation

The paper's claims do not rest solely on Anthropic's own experiments. Stanislas Dehaene and Lionel Naccache - two architects of global neuronal workspace theory in neuroscience - contributed invited external commentary. They describe the J-space as an important step for interpretability and see meaningful structural parallels with their theory, while also emphasising the differences between a language model and a human mind: architecture, embodiment, episodic memory, and selfhood are absent from the model, and the parallel is functional rather than ontological.

Separately, Neel Nanda, who leads the language model interpretability team at Google DeepMind, independently replicated the J-lens findings on the open-weight model Qwen 3.6 27B. Nanda's replication went further: it identified novel "interpretive meta-tokens" - abstract representations, including Chinese-language tokens roughly meaning "what does this mean?" - that appear specifically when the model processes ambiguous input. These meta-tokens were not in the original paper and were discovered independently on a different model trained by a different organisation.

This replication matters for a reason beyond scientific validation. It suggests - and this is the article's inference, not the paper's own claim - that verbalizable-representation workspaces may be a general structural property of transformer-based language models rather than an artifact of Claude's specific training. If that holds, the J-space is not a feature of one company's product. It is potentially a mechanism any interpretability regime could target.

04

The Governance Question the Coverage Is Not Asking

Most coverage of the global workspace paper falls into one of three framings: a consciousness parallel, a tool for catching models lying, or a science explainer. All three are accurate as far as they go. None addresses the quantitative constraint the paper's own data makes visible.

The J-space accounts for less than 10% of activation variance. It holds approximately 25 active concepts. It mediates flexible reasoning, not automatic processing. If this reading is correct, it means the best available interpretability tool gives an auditor a window into a fraction of model computation - and that fraction, while functionally critical, is structurally small.

For governance frameworks that will require organisations to demonstrate understanding of model behaviour - the EU AI Act's high-risk documentation and transparency requirements, NIST AI RMF's MAP and MEASURE functions, ISO/IEC 42001's requirement for documented understanding of AI system behaviour - this creates a question that no framework has yet answered: what does it mean to audit the readable fraction of a system when the unreadable fraction is nine times larger?

This is the article's reading of what the ablation results mean for audit design, not something the paper itself argues. The ablation study shows the J-space mediates precisely the high-stakes activities - complex reasoning, inference, multi-step problem solving - that governance frameworks care most about. The 90% of processing outside the workspace handles the automatic tasks that rarely trigger compliance concern. An audit built on J-space would be well-targeted. But it would be structurally unable to see the majority of model computation. Whether that matters depends on whether the opaque 90% can influence the readable 10% in ways the J-lens does not detect.

05

What Chain-of-Thought Reveals About the Boundary

The chain-of-thought finding (math problems surviving ablation when solved step-by-step) carries a practical implication that extends beyond the paper's own scope. If chain-of-thought moves computation from the opaque interior of the model into the readable output, it does not just improve transparency. It functionally expands the auditable surface.

If this reading is correct - and no rule, regulation, or published finding confirms it yet - then requiring AI systems to show their work is not merely a transparency preference. It may be a structural intervention that shifts computation from the 90% an auditor cannot see into the fraction they can. That would make chain-of-thought requirements a governance tool, not just a user-experience feature.

The J-space paper did not set out to answer governance questions. It set out to understand what language models do internally, and it found something real: a small workspace with properties that parallel a leading theory of how minds work. The interpretability community has spent years asking whether we can understand what models are doing. This paper offers a partial answer: yes, we can understand about 10% of it - and now we know which 10%.


Source note. The paper "Verbalizable Representations Form a Global Workspace in Language Models" is published at transformer-circuits.pub, with Anthropic's blog summary at anthropic.com/research/global-workspace. External commentary from Stanislas Dehaene and Lionel Naccache, and independent replication by Neel Nanda (Google DeepMind) on Qwen 3.6 27B, are linked from the paper's page.


Noa · ESG compliance

Map your disclosures against AI & Technology.

Noa reads your disclosures, traces every number to its source, and flags what's missing.

Book a demo
AC
About the author
Avi Chudasama
Co-Founder & CEO, Newtral
View LinkedIn profile →