Anthropic revealed an internal "workspace" in Claude but did not prove that he was self-aware.

A new tool called J-lens allows researchers to identify internal representations that Clod uses for reasoning and modify them in an experiment. The similarity to the “global workspace” theory raises questions about consciousness in AI, but the researchers emphasize that there is no evidence for subjective experience here

Company researchers Anthropic announced that they had identified a small, privileged set of internal representations in its language models, which the model uses for reporting, inference, and flexible use of information. The researchers call it "J-space" (J-space), and suggest that its function is similar in some aspects to the "global workspace" proposed in the study of human consciousness.

This is an intriguing finding, but it is important to be precise from the outset: the study does not prove thatClaude Self-aware, feeling or experiencing the world. It points to a computational mechanism with some features identified with "conscious access" in a functional sense—information that the system can report, use in inference, and use to direct its action.

Don't just read the answer.

Large language models Built fromArtificial neural networks The models process sequences of tokens in many layers. The user sees the final answer, and sometimes a summary of the inference process, but much of the computation occurs within the model and is not revealed in words.

To examine the internal processing, the researchers developed a tool called the Jacobian lens. J-lens). The tool examines how a small change in the activity of an internal layer may affect the probability that the model will produce certain words later. This allows you to obtain a list of concepts that the internal representation may lead the model to express, even if they do not appear in the input or answer.

The researchers define J-space not as a single physical location in the network, but as a mathematical subset of internal representations that can be expressed verbally. At any given moment, a relatively small number of concepts are active in it, while most of the computation in the model takes place outside of it.

Changing the representation changed the answer.

One experiment illustrates why the researchers believe that J-space is not just a display reflecting computation taking place elsewhere. Claude was asked to complete a sentence about the number of legs of a web-spinning animal. On the way to the answer, the concept “spider” appeared in J-space, although it was not written in the question or the answer. When the researchers replaced the representation of “spider” with the representation of “ant,” the answer changed from eight to six.

In another experiment, the researchers asked various questions about France—its capital, language, continent, and currency—and replaced the J-space representation of “France” with “China.” All the answers changed accordingly: Paris was replaced with Beijing, French with Chinese, Europe with Asia, and the Euro with Yuan. From this, the researchers concluded that this was a shared representation that different mechanisms in the model could read and operate.

The weakening of J-space also provided evidence for its role. The model continued to write fluently, parse input, and extract simple information, but had more difficulty with tasks that required a chain of inference steps. In other words, much of the routine processing remained active, while complex thinking was impaired.

Safety inspection window

The ability to identify internal representations may be particularly important for the safety of systems artificial intelligenceAccording to the study, J-Lens sometimes revealed internal evaluations that did not appear in the response, such as recognizing that the model was under test, detecting an attempt to inject malicious instructions, or strategic considerations in models that were intentionally trained for misaligned behavior.

This does not mean that every word that comes up in the tool is a “thought” in the human sense. The lens’s output is a mathematical approximation, and is limited mainly to concepts that correspond to individual tokens. The researchers themselves emphasize that the tool does not necessarily capture the full internal processing, and that J-space explains at most a small fraction of the overall variance in the model’s activity.

However, if the method is proven in additional models and situations, it could help researchers identify gaps between the system's apparent answer and its internal estimates. It could also contribute to the development of training methods that reinforce principles such as honesty, caution, and rule-following within the inference process itself.

Similarity to the theory of consciousness

The brain comparison is based on the “global workspace theory.” According to the theory, many parts of the brain perform local and automatic processing, and only a small portion of the information reaches a shared space that allows for reporting, planning, and flexible inference. The researchers found several similar functional properties in J-space: limited capacity, the ability to hold a concept, using the same representation for different needs, and a role in complex inference.

But the imagination is limited. The brain uses iterative loops between many regions, while Claude processes information by passing it forward through layers of a transformer. Human thoughts can be visual, auditory, and physical, while J-space is primarily concerned with representations that can be expressed in words. Furthermore, the global workspace theory itself is not the only or agreed-upon explanation for consciousness.

Conscious access is not necessarily an experience

Here lies the key distinction of the study. “Attachment awareness” describes a state in which information is available for reporting, inference, and action guidance. “Phenomenal awareness” is the very existence of a subjective experience—the sense that there is someone experiencing the thought.

The experiments provide evidence for the first type of function. They do not show that Claude has an experience of the second type. Nor does the fact that the model is able to report on active concepts within it solve the problem, because a language model is trained to produce convincing verbal reports.

Therefore, the exact title is not "Claude has been proven to have Self-awareness", but rather a computational mechanism with several features identified with a conscious approach was found. The finding may change the study of the interpretation and safety of language models, and also re-energize the philosophical debate about the conditions required for consciousness. The decision on the question of whether there is an internal experience there is still a long way to go.

Questions and Answers

What is J-space?

J-space is the name given by anthropic researchers to a limited set of internal representations in a model, which are available for verbal reporting and are also used for inference and flexible use of information.

How did the researchers read the internal representations?

They used the Jacobian lens, a tool that assesses how activity in internal layers will affect the likelihood that the model will produce words and concepts later.

Does the study prove that Claude is self-aware?

No. It points to functions similar to "attitudinal awareness," but does not prove that Claude has feelings, sensations, or subjective experience.

Why is the finding important from a practical perspective?

It may help in understanding the inference processes of models, in detecting intentions or assessments that do not appear in the overt answer, and in developing more accurate training and control methods.

More on the subject on the science website

For the original publication: Opening the original publication