The Essence of Contextual Understanding in Theory of Mind: A Study on Question Answering with Story Characters
Chulun Zhou, Qiujing Wang, Mo Yu, Xiaoqian Yue, Rui Lu, Jiangnan Li, Yifan Zhou, Shunchi Zhang, Jie Zhou, Wai Lam
摘要
Theory-of-Mind (ToM) is a fundamental psychological capability that allows humans to understand and interpret the mental states of others. Humans infer others' thoughts by integrating causal cues and indirect clues from broad contextual information, often derived from past interactions. In other words, human ToM heavily relies on the understanding about the backgrounds and life stories of others. Unfortunately, this aspect is largely overlooked in existing benchmarks for evaluating machines' ToM capabilities, due to their usage of short narratives without global context, especially personal background of characters. In this paper, we verify the importance of comprehensive contextual understanding about personal backgrounds in ToM and assess the performance of LLMs in such complex scenarios. To achieve this, we introduce CHARTOM-QA benchmark, comprising 1,035 ToM questions based on characters from classic novels. Our human study reveals a significant disparity in performance: the same group of educated participants performs dramatically better when they have read the novels compared to when they have not. In parallel, our experiments on state-of-the-art LLMs, including the very recent o1 and DeepSeek-R1 models, show that LLMs still perform notably worse than humans, despite that they have seen these stories during pre-training. This highlights the limitations of current LLMs in capturing the nuanced contextual information required for ToM reasoning. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- HGMem: Hypergraph-based Working Memory to Improve Multi-step RAG for Long-Context Complex Relational ModelingChulun Zhou, Chunkang Zhang, Guoxin Yu, Fandong Meng 等ICML 2026 · 被引用 4 次
- MemGraphRAG: Memory-based Multi-Agent System for Graph Retrieval-Augmented GenerationChuanjie Wu, Zhishang Xiang, Yunbo Tang, Zerui Chen 等KDD 2026 · 被引用 1 次
- MindZero: Learning Online Mental Reasoning With Zero AnnotationsShunchi Zhang, Jin Lu, Chuanyang Jin, Yichao Zhou 等ICML 2026 · 被引用 1 次
它引用的顶会 Paper9
- Fantastic Questions and Where to Find Them: FairytaleQA - An Authentic Dataset for Narrative ComprehensionYing Xu, Dakuo Wang, Mo Yu, Daniel Ritchie 等ACL 2022 · 被引用 131 次
- Few-Shot Character Understanding in Movies as an Assessment to Meta-Learning of Theory-of-MindMo Yu, Qiujing Wang, Shunchi Zhang, Yisi Sang 等ICML 2024 · 被引用 22 次
- RELiC: Retrieving Evidence for Literary ClaimsKatherine Thai, Yapei Chang, Kalpesh Krishna, Mohit IyyerACL 2022 · 被引用 22 次
- FANToM: A Benchmark for Stress-testing Machine Theory of Mind in InteractionsHyunwoo Kim, Melanie Sclar, Xuhui Zhou, Ronan Le Bras 等EMNLP 2023 · 被引用 21 次
- ToMBench: Benchmarking Theory of Mind in Large Language ModelsZhuang Chen, Jincenzi Wu, Jinfeng Zhou, Bosi Wen 等ACL 2024 · 被引用 6 次
相关 Paper
- OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language ModelsHainiu Xu, Runcong Zhao, Lixing Zhu, Jinhua Du 等ACL 2024
- Theory of Mind in Large Language Models: Assessment and EnhancementRuirui Chen, Weifeng Jiang, Chengwei Qin, Cheston TanACL 2025
- XToM: Exploring the Multilingual Theory of Mind for Large Language ModelsChunkit Chan, Yauwai Yim, Hongchuan Zeng, Zhiying Zou 等ACL 2026
- Perceptions to Beliefs: Exploring Precursory Inferences for Theory of Mind in Large Language ModelsChani Jung, Dongkwan Kim, Jiho Jin, Jiseon Kim 等EMNLP 2024 · 被引用 2 次
- Minding Language Models' (Lack of) Theory of Mind: A Plug-and-Play Multi-Character Belief TrackerMelanie Sclar, Sachin Kumar, Peter West, Alane Suhr 等ACL 2023 · 被引用 21 次
