HeadlinesBriefing favicon HeadlinesBriefing.com

LLMs Show Introspective Awareness of Internal States

Hacker News •
×

We investigate whether large language models can introspect on their internal states. It is difficult to distinguish genuine introspection from confabulations through conversation alone. We address this by injecting representations of known concepts into a model's activations and measuring the influence on self-reported states.

We find that models can notice the presence of injected concepts and accurately identify them. Models demonstrate some ability to recall prior internal representations and distinguish them from raw text inputs. Strikingly, some models can use their ability to recall prior intentions to distinguish their own outputs from artificial prefills.

In all experiments, Claude Opus 4 and 4.1 generally demonstrate the greatest introspective awareness; however, trends across models are complex and sensitive to post-training strategies. Finally, we explore whether models can explicitly control their internal representations, finding that models can modulate their activations when instructed or incentivized to "think about" a concept. Overall, current language models possess some functional introspective awareness of their internal states, but this capacity is highly unreliable and context-dependent.

It may continue to develop with further improvements. Submission history: From Jack Lindsey [view email] [v1]Mon, 5 Jan 2026 06:47:41 UTC.