MovieGraph-ToM: Evaluating Long-Range Theory of Mind in Large Language Models via Implicit Social-Causal Graphs
Tingjiang Wei, Qin Ni, Rong Gao, Yingying Wang, Liang He
Abstract
The capacity for social reasoning, particularly Theory of Mind (ToM), is a foundational prerequisite for aligning Large Language Models (LLMs) with human values. However, current evaluations are predominantly confined to simplistic, short-text scenarios, obscuring their true capabilities and potential failure modes in complex, long-range social dynamics. To address this deficit, we introduce MovieGraph-ToM, a large-scale benchmark for evaluating long-range ToM and social cognition within extended, multimodal narratives. We employ a "scaffold-and-probe" methodology, and we construct a ground-truth Social-Causal Graph offline, which maps the narrative's latent mental states and causal chains. During evaluation, the model is denied access to this graph and must reason directly from raw multimodal inputs. This decoupling forces genuine inference over superficial pattern matching. Reasoning is probed via a hierarchical questioning framework designed to differentiate spontaneous understanding from logical robustness. Our empirical results reveal systematic vulnerabilities in even state-of-the-art models. We identify a critical multiple-choice pitfall, where accuracy plummets against well-crafted distractors, and a stark "generative-discriminative divide," where models fail to construct coherent explanations for answers they correctly identify. These findings highlight a latent risk, as models that feign comprehension could lead to unpredictable and misaligned behaviors. MovieGraph-ToM thus offers a rigorous platform for assessing and advancing the robust social intelligence required for safely aligned AI systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 67ff8bae-48fe-4432-bb71-8f2a411019e5Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou et al.ICLR 2024 · 1,197 citations
- SimpleToM: Exposing the Gap between Explicit ToM Inference and Implicit ToM Application in LLMsYuling Gu, Oyvind Tafjord, Hyunwoo Kim, Jared Moore et al.ICLR 2026 · 39 citations
- AgentGym: Evaluating and Training Large Language Model-based Agents across Diverse EnvironmentsZhiheng Xi, Yiwen Ding, Wenxiang Chen, Boyang Hong et al.ACL 2025 · 20 citations
- ToMATO: Verbalizing the Mental States of Role-Playing LLMs for Benchmarking Theory of MindKazutoshi Shinoda, Nobukatsu Hojo, Kyosuke Nishida, Saki Mizuno et al.AAAI 2025 · 10 citations
Related papers
- Theory of Mind in Large Language Models: Assessment and EnhancementRuirui Chen, Weifeng Jiang, Chengwei Qin, Cheston TanACL 2025
- ToMBench: Benchmarking Theory of Mind in Large Language ModelsZhuang Chen, Jincenzi Wu, Jinfeng Zhou, Bosi Wen et al.ACL 2024 · 6 citations
- Explore Theory of Mind: program-guided adversarial data generation for theory of mind reasoningMelanie Sclar, Jane Dwivedi-Yu, Maryam Fazel-Zarandi, Yulia Tsvetkov et al.ICLR 2025
- HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized BenchmarksTing Zhou, Daoyuan Chen, Qirui Jiao, Bolin Ding et al.CVPR 2026
- GroupToM-Bench: Benchmarking Group Theory of Mind and Nonlinear Social Emergence in MLLMsWeidong Tang, Jierui Li, Yueling Hou, Zihan Mei et al.ACL 2026
