Do language models have coherent mental models of everyday things?
Yuling Gu, Bhavana Dalvi Mishra, Peter Clark
摘要
When people think of everyday things like an egg, they typically have a mental image associated with it. This allows them to correctly judge, for example, that "the yolk surrounds the shell" is a false statement. Do language models similarly have a coherent picture of such everyday things? To investigate this, we propose a benchmark dataset consisting of 100 everyday things, their parts, and the relationships between these parts, expressed as 11,720 "X relation Y?" true/false questions. Using these questions as probes, we observe that state-ofthe-art pre-trained language models (LMs) like GPT-3 and Macaw have fragments of knowledge about these everyday things, but do not have fully coherent "parts mental models" (54-59% accurate, 19-43% conditional constraint violation). We propose an extension where we add a constraint satisfaction layer on top of the LM's raw predictions to apply commonsense constraints. As well as removing inconsistencies, we find that this also significantly improves accuracy (by 16-20%), suggesting how the incoherence of the LM's pictures of everyday things can be significantly reduced. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- SimpleToM: Exposing the Gap between Explicit ToM Inference and Implicit ToM Application in LLMsYuling Gu, Oyvind Tafjord, Hyunwoo Kim, Jared Moore 等ICLR 2026 · 被引用 39 次
- Language Models with RationalityNora Kassner, Oyvind Tafjord, Ashish Sabharwal, Kyle Richardson 等EMNLP 2023 · 被引用 7 次
- EpiK-Eval: Evaluation for Language Models as Epistemic ModelsGabriele Prato, Jerry Huang, Prasanna Parthasarathi, Shagun Sodhani 等EMNLP 2023
- Aligning with Logic: Measuring, Evaluating and Improving Logical Preference Consistency in Large Language ModelsYinhong Liu, Zhijiang Guo, Tianya Liang, Ehsan Shareghi 等ICML 2025
它引用的顶会 Paper5
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- Improving Coherence and Consistency in Neural Sequence Models with Dual-System, Neuro-Symbolic ReasoningMaxwell I. Nye, Michael Henry Tessler, Joshua B. Tenenbaum, Brenden M. LakeNeurIPS 2021 · 被引用 151 次
- PTR: A Benchmark for Part-based Conceptual, Relational, and Physical ReasoningYining Hong, Li Yi, Josh Tenenbaum, Antonio Torralba 等NeurIPS 2021 · 被引用 46 次
- Abstract Visual Reasoning with Tangram ShapesAnya Ji, Noriyuki Kojima, Noah Rush, Alane Suhr 等EMNLP 2022 · 被引用 18 次
- BeliefBank: Adding Memory to a Pre-Trained Language Model for a Systematic Notion of BeliefNora Kassner, Oyvind Tafjord, Hinrich Schütze, Peter ClarkEMNLP 2021 · 被引用 2 次
相关 Paper
- COPEN: Probing Conceptual Knowledge in Pre-trained Language ModelsHao Peng, Xiaozhi Wang, Shengding Hu, Hailong Jin 等EMNLP 2022 · 被引用 16 次
- FOCUS: Evaluating Pre-trained Vision-Language Models on Underspecification ReasoningKankan Zhou, Eason Lai, Kyriakos Mouratidis, Jing JiangACL 2025
- TIMEDIAL: Temporal Commonsense Reasoning in DialogLianhui Qin, Aditya Gupta, Shyam Upadhyay, Luheng He 等ACL 2021
- Mapping Language Models to Grounded Conceptual SpacesRoma Patel, Ellie PavlickICLR 2022 · 被引用 197 次
- The Sensitivity of Language Models and Humans to Winograd Schema PerturbationsMostafa Abdou, Vinit Ravishankar, Maria Barrett, Yonatan Belinkov 等ACL 2020 · 被引用 1 次
