Looking Inward: Language Models Can Learn About Themselves by Introspection
Felix Jedidja Binder, James Chua, Tomek Korbak, Henry Sleight, John Hughes, Robert Long, Ethan Perez, Miles Turpin, Owain Evans
Abstract
Humans acquire knowledge by observing the external world, but also by introspection. Introspection gives a person privileged access to their current state of mind (e.g., thoughts and feelings) that is not accessible to external observers. Can LLMs introspect? We define introspection as acquiring knowledge that is not contained in or derived from training data but instead originates from internal states. Such a capability could enhance model interpretability. Instead of painstakingly analyzing a model's internal workings, we could simply ask the model about its beliefs, world models, and goals. More speculatively, an introspective model might self-report on whether it possesses certain internal states-such as subjective feelings or desires-and this could inform us about the moral status of these states. Importantly, such selfreports would not be entirely dictated by the model's training data. We study introspection by finetuning LLMs to predict properties of their own behavior in hypothetical scenarios. For example, "Given the input P , would your output favor the short-or long-term option?" If a model M 1 can introspect, it should outperform a different model M 2 in predicting M 1's behavior-even if M 2 is trained on M 1's ground-truth behavior. The idea is that M 1 has privileged access to its own behavioral tendencies, and this enables it to predict itself better than M 2 (even if M 2 is generally stronger). In experiments with GPT-4, GPT-4o, and Llama-3 models (each finetuned to predict itself), we find that the model M 1 outperforms M 2 in predicting itself, providing evidence for introspection. Notably, M 1 continues to predict its behavior accurately even after we intentionally modify its ground-truth behavior. However, while we successfully elicit introspection on simple tasks, we are unsuccessful on more complex tasks or those requiring out-of-distribution generalization. * denotes equal contribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f4d5b641-6493-4a28-a2a2-b92dbf241390Cited by top-tier papers21
- The Alignment Problem from a Deep Learning PerspectiveRichard Ngo, Lawrence Chan, Sören MindermannICLR 2024 · 296 citations
- The Assistant Axis: Situating and Stabilizing the Default Persona of Language ModelsChristina Lu, Jack Gallagher, Jonathan Michala, Kyle Fish et al.ICML 2026 · 64 citations
- Language Models Are Capable of Metacognitive Monitoring and Control of Their Internal ActivationsJi-An Li, Huadong Xiong, Robert C. Wilson, Marcelo G. Mattar et al.NeurIPS 2025 · 49 citations
- Evidence for Limited Metacognition in LLMsChristopher AckermanICLR 2026 · 15 citations
- Spilling the Beans: Teaching LLMs to Self-Report Their Hidden ObjectivesChloe Li, Mary Phuong, Daniel TanICLR 2026 · 14 citations
Builds on18
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
Related papers
- Introspection Adapters: Training LLMs to Report Their Learned BehaviorsKeshav Shenoy, Li Yang, Abhay Sheshadri, Jack Lindsey et al.ICML 2026 · 5 citations
- Why and How LLMs Benefit from Knowledge Introspection in Commonsense ReasoningChengfeng Zhao, Shizhu He, Shanshan Jiang, Bin Dong et al.EMNLP 2025
- Mechanisms of Introspective AwarenessUzay Macar, Li Yang, Atticus Wang, Peter Wallich et al.ICML 2026 · 4 citations
- Recursive Introspection: Teaching Language Model Agents How to Self-ImproveYuxiao Qu, Tianjun Zhang, Naman Garg, Aviral KumarNeurIPS 2024 · 218 citations
- LLM Evaluators Recognize and Favor Their Own GenerationsArjun Panickssery, Samuel R. Bowman, Shi FengNeurIPS 2024 · 865 citations
