Language Models Represent Beliefs of Self and Others
Wentao Zhu, Zhining Zhang, Yizhou Wang
摘要
Understanding and attributing mental states, known as Theory of Mind (ToM), emerges as a fundamental capability for human social reasoning. While Large Language Models (LLMs) appear to possess certain ToM abilities, the mechanisms underlying these capabilities remain elusive. In this study, we discover that it is possible to linearly decode the belief status from the perspectives of various agents through neural activations of language models, indicating the existence of internal representations of self and others' beliefs. By manipulating these representations, we observe dramatic changes in the models' ToM performance, underscoring their pivotal role in the social reasoning process. Additionally, our findings extend to diverse social reasoning tasks that involve different causal inference patterns, suggesting the potential generalizability of these representations. *
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- From Flat to Hierarchical: Extracting Sparse Representations with Matching PursuitValérie Costa, Thomas Fel, Ekdeep Singh Lubana, Bahareh Tolooshams 等NeurIPS 2025 · 被引用 54 次
- Language Models Use Lookbacks to Track BeliefsNikhil Prakash, Natalie Shapira, Arnab Sen Sharma, Christoph Riedl 等ICLR 2026 · 被引用 42 次
- Toward Consistent World Models with Multi-Token Prediction and Latent Semantic EnhancementQimin Zhong, Hao Liao, Haiming Qin, Mingyang Zhou 等ACL 2026 · 被引用 3 次
- The Energy Loss Phenomenon in RLHF: A New Perspective on Mitigating Reward HackingYuchun Miao, Sen Zhang, Liang Ding, Yuqi Zhang 等ICML 2025
- Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation SteeringEric Bigelow, Daniel Wurgaft, YingQiao Wang, Noah Goodman 等ICML 2026
它引用的顶会 Paper9
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister 等NeurIPS 2023 · 被引用 1,549 次
- The Linear Representation Hypothesis and the Geometry of Large Language ModelsKiho Park, Yo Joong Choe, Victor VeitchICML 2024 · 被引用 461 次
- Language Models Represent Space and TimeWes Gurnee, Max TegmarkICLR 2024 · 被引用 303 次
- The Alignment Problem from a Deep Learning PerspectiveRichard Ngo, Lawrence Chan, Sören MindermannICLR 2024 · 被引用 296 次
相关 Paper
- Theory of Mind for Multi-Agent Collaboration via Large Language ModelsHuao Li, Yu Quan Chong, Simon Stepputtis, Joseph Campbell 等EMNLP 2023 · 被引用 57 次
- AutoToM: Scaling Model-based Mental Inference via Automated Agent ModelingZhining Zhang, Chuanyang Jin, Mung Yao Jia, Shunchi Zhang 等NeurIPS 2025 · 被引用 30 次
- Minding Language Models' (Lack of) Theory of Mind: A Plug-and-Play Multi-Character Belief TrackerMelanie Sclar, Sachin Kumar, Peter West, Alane Suhr 等ACL 2023 · 被引用 21 次
- Perceptions to Beliefs: Exploring Precursory Inferences for Theory of Mind in Large Language ModelsChani Jung, Dongkwan Kim, Jiho Jin, Jiseon Kim 等EMNLP 2024 · 被引用 2 次
- MetaMind: Modeling Human Social Thoughts with Metacognitive Multi-Agent SystemsXuanming Zhang, Yuxuan Chen, Samuel (Min-Hsuan) Yeh, Sharon LiNeurIPS 2025 · 被引用 14 次
