OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models
Hainiu Xu, Runcong Zhao, Lixing Zhu, Jinhua Du, Yulan He
Abstract
Neural Theory-of-Mind (N-ToM), machine's ability to understand and keep track of the mental states of others, is pivotal in developing socially intelligent agents. However, prevalent N-ToM benchmarks have several shortcomings, including the presence of ambiguous and artificial narratives, absence of personality traits and preferences, a lack of questions addressing characters' psychological mental states, and limited diversity in the questions posed. In response to these issues, we construct Open-ToM, a new benchmark for assessing N-ToM with (1) longer and clearer narrative stories, (2) characters with explicit personality traits, (3) actions that are triggered by character intentions, and (4) questions designed to challenge LLMs' capabilities of modeling characters' mental states of both the physical and psychological world. Using OpenToM, we reveal that state-of-the-art LLMs thrive at modeling certain aspects of mental states in the physical world but fall short when tracking characters' mental states in the psychological world. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1e16d173-2125-462a-a3c2-a458fa2d7a56Cited by top-tier papers29
- MuMA-ToM: Multi-modal Multi-Agent Theory of MindHaojun Shi, Suyu Ye, Xinyu Fang, Chuanyang Jin et al.AAAI 2025 · 48 citations
- Language Models Use Lookbacks to Track BeliefsNikhil Prakash, Natalie Shapira, Arnab Sen Sharma, Christoph Riedl et al.ICLR 2026 · 42 citations
- SimpleToM: Exposing the Gap between Explicit ToM Inference and Implicit ToM Application in LLMsYuling Gu, Oyvind Tafjord, Hyunwoo Kim, Jared Moore et al.ICLR 2026 · 39 citations
- AutoToM: Scaling Model-based Mental Inference via Automated Agent ModelingZhining Zhang, Chuanyang Jin, Mung Yao Jia, Shunchi Zhang et al.NeurIPS 2025 · 30 citations
- Autonomous Agents for Collaborative Task under Information AsymmetryWei Liu, Chenxi Wang, Yifei Wang, Zihao Xie et al.NeurIPS 2024 · 17 citations
Builds on13
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- PAL: Program-aided Language ModelsLuyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon et al.ICML 2023 · 700 citations
Related papers
- The Essence of Contextual Understanding in Theory of Mind: A Study on Question Answering with Story CharactersChulun Zhou, Qiujing Wang, Mo Yu, Xiaoqian Yue et al.ACL 2025 · 9 citations
- ToMBench: Benchmarking Theory of Mind in Large Language ModelsZhuang Chen, Jincenzi Wu, Jinfeng Zhou, Bosi Wen et al.ACL 2024 · 6 citations
- ToMATO: Verbalizing the Mental States of Role-Playing LLMs for Benchmarking Theory of MindKazutoshi Shinoda, Nobukatsu Hojo, Kyosuke Nishida, Saki Mizuno et al.AAAI 2025 · 10 citations
- Theory of Mind in Large Language Models: Assessment and EnhancementRuirui Chen, Weifeng Jiang, Chengwei Qin, Cheston TanACL 2025
- Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human StatesYang Xiao, Jiashuo Wang, Qiancheng Xu, Changhe Song et al.ACL 2025 · 12 citations
