Looking Inward: Language Models Can Learn About Themselves by Introspection
Felix Jedidja Binder, James Chua, Tomek Korbak, Henry Sleight, John Hughes, Robert Long, Ethan Perez, Miles Turpin, Owain Evans
摘要
Humans acquire knowledge by observing the external world, but also by introspection. Introspection gives a person privileged access to their current state of mind (e.g., thoughts and feelings) that is not accessible to external observers. Can LLMs introspect? We define introspection as acquiring knowledge that is not contained in or derived from training data but instead originates from internal states. Such a capability could enhance model interpretability. Instead of painstakingly analyzing a model's internal workings, we could simply ask the model about its beliefs, world models, and goals. More speculatively, an introspective model might self-report on whether it possesses certain internal states-such as subjective feelings or desires-and this could inform us about the moral status of these states. Importantly, such selfreports would not be entirely dictated by the model's training data. We study introspection by finetuning LLMs to predict properties of their own behavior in hypothetical scenarios. For example, "Given the input P , would your output favor the short-or long-term option?" If a model M 1 can introspect, it should outperform a different model M 2 in predicting M 1's behavior-even if M 2 is trained on M 1's ground-truth behavior. The idea is that M 1 has privileged access to its own behavioral tendencies, and this enables it to predict itself better than M 2 (even if M 2 is generally stronger). In experiments with GPT-4, GPT-4o, and Llama-3 models (each finetuned to predict itself), we find that the model M 1 outperforms M 2 in predicting itself, providing evidence for introspection. Notably, M 1 continues to predict its behavior accurately even after we intentionally modify its ground-truth behavior. However, while we successfully elicit introspection on simple tasks, we are unsuccessful on more complex tasks or those requiring out-of-distribution generalization. * denotes equal contribution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- The Alignment Problem from a Deep Learning PerspectiveRichard Ngo, Lawrence Chan, Sören MindermannICLR 2024 · 被引用 296 次
- The Assistant Axis: Situating and Stabilizing the Default Persona of Language ModelsChristina Lu, Jack Gallagher, Jonathan Michala, Kyle Fish 等ICML 2026 · 被引用 64 次
- Language Models Are Capable of Metacognitive Monitoring and Control of Their Internal ActivationsJi-An Li, Huadong Xiong, Robert C. Wilson, Marcelo G. Mattar 等NeurIPS 2025 · 被引用 49 次
- Evidence for Limited Metacognition in LLMsChristopher AckermanICLR 2026 · 被引用 15 次
- Spilling the Beans: Teaching LLMs to Self-Report Their Hidden ObjectivesChloe Li, Mary Phuong, Daniel TanICLR 2026 · 被引用 14 次
它引用的顶会 Paper18
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos 等ICML 2024 · 被引用 1,212 次
相关 Paper
- Introspection Adapters: Training LLMs to Report Their Learned BehaviorsKeshav Shenoy, Li Yang, Abhay Sheshadri, Jack Lindsey 等ICML 2026 · 被引用 5 次
- Why and How LLMs Benefit from Knowledge Introspection in Commonsense ReasoningChengfeng Zhao, Shizhu He, Shanshan Jiang, Bin Dong 等EMNLP 2025
- Mechanisms of Introspective AwarenessUzay Macar, Li Yang, Atticus Wang, Peter Wallich 等ICML 2026 · 被引用 4 次
- Recursive Introspection: Teaching Language Model Agents How to Self-ImproveYuxiao Qu, Tianjun Zhang, Naman Garg, Aviral KumarNeurIPS 2024 · 被引用 218 次
- LLM Evaluators Recognize and Favor Their Own GenerationsArjun Panickssery, Samuel R. Bowman, Shi FengNeurIPS 2024 · 被引用 865 次
