ToMATO: Verbalizing the Mental States of Role-Playing LLMs for Benchmarking Theory of Mind
Kazutoshi Shinoda, Nobukatsu Hojo, Kyosuke Nishida, Saki Mizuno, Keita Suzuki, Ryo Masumura, Hiroaki Sugiyama, Kuniko Saito
Abstract
Existing Theory of Mind (ToM) benchmarks diverge from real-world scenarios in three aspects: 1) they assess a limited range of mental states such as beliefs, 2) false beliefs are not comprehensively explored, and 3) the diverse personality traits of characters are overlooked. To address these challenges, we introduce ToMATO, a new ToM benchmark formulated as multiple-choice QA over conversations. ToMATO is generated via LLM-LLM conversations featuring information asymmetry. By employing a prompting method that requires role-playing LLMs to verbalize their thoughts before each utterance, we capture both first- and second-order mental states across five categories: belief, intention, desire, emotion, and knowledge. These verbalized thoughts serve as answers to questions designed to assess the mental states of characters within conversations. Furthermore, the information asymmetry introduced by hiding thoughts from others induces the generation of false beliefs about various mental states. Assigning distinct personality traits to LLMs further diversifies both utterances and thoughts. ToMATO consists of 5.4k questions, 753 conversations, and 15 personality trait patterns. Our analysis shows that this dataset construction approach frequently generates false beliefs due to the information asymmetry between role-playing LLMs, and effectively reflects diverse personalities. We evaluate nine LLMs on ToMATO and find that even GPT-4o mini lags behind human performance, especially in understanding false beliefs, and lacks robustness to various personality traits.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d74b4d8f-81bc-4e2f-a1f7-ea66657a3916Cited by top-tier papers3
- GroupToM-Bench: Benchmarking Group Theory of Mind and Nonlinear Social Emergence in MLLMsWeidong Tang, Jierui Li, Yueling Hou, Zihan Mei et al.ACL 2026
- Tracing Belief-Driven Thoughts with Theory-of-Mind Agents: An Opinion Analysis FrameworkJintao Wen, Yunfeng Ning, Hankun Kang, Xin Miao et al.WWW 2026
- MovieGraph-ToM: Evaluating Long-Range Theory of Mind in Large Language Models via Implicit Social-Causal GraphsTingjiang Wei, Qin Ni, Rong Gao, Yingying Wang et al.AAAI 2026
Builds on16
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- ReClor: A Reading Comprehension Dataset Requiring Logical ReasoningWeihao Yu, Zihang Jiang, Yanfei Dong, Jiashi FengICLR 2020 · 325 citations
- Evaluating and Inducing Personality in Pre-trained Language ModelsGuangyuan Jiang, Manjie Xu, Song-Chun Zhu, Wenjuan Han et al.NeurIPS 2023 · 192 citations
Related papers
- FANToM: A Benchmark for Stress-testing Machine Theory of Mind in InteractionsHyunwoo Kim, Melanie Sclar, Xuhui Zhou, Ronan Le Bras et al.EMNLP 2023 · 21 citations
- ToMBench: Benchmarking Theory of Mind in Large Language ModelsZhuang Chen, Jincenzi Wu, Jinfeng Zhou, Bosi Wen et al.ACL 2024 · 6 citations
- OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language ModelsHainiu Xu, Runcong Zhao, Lixing Zhu, Jinhua Du et al.ACL 2024
- Few-Shot Character Understanding in Movies as an Assessment to Meta-Learning of Theory-of-MindMo Yu, Qiujing Wang, Shunchi Zhang, Yisi Sang et al.ICML 2024 · 22 citations
- CogToM: A Comprehensive Theory of Mind Benchmark inspired by Human Cognition for Large Language ModelsHaibo Tong, Zeyang Yue, Feifei Zhao, Erliang Lin et al.ACL 2026
