FANToM: A Benchmark for Stress-testing Machine Theory of Mind in Interactions
Hyunwoo Kim, Melanie Sclar, Xuhui Zhou, Ronan Le Bras, Gunhee Kim, Yejin Choi, Maarten Sap
摘要
Theory of mind (ToM) evaluations currently focus on testing models using passive narratives that inherently lack interactivity. We introduce FANTOM, a new benchmark designed to stress-test ToM within information-asymmetric conversational contexts via question answering. Our benchmark draws upon important theoretical requisites from psychology and necessary empirical considerations when evaluating large language models (LLMs). In particular, we formulate multiple types of questions that demand the same underlying reasoning to identify illusory or false sense of ToM capabilities in LLMs. We show that FANTOM is challenging for state-of-the-art LLMs, which perform significantly worse than humans even with chainof-thought reasoning or fine-tuning. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper35
- Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity TheoryNiloofar Mireshghallah, Hyunwoo Kim, Xuhui Zhou, Yulia Tsvetkov 等ICLR 2024 · 被引用 198 次
- MuMA-ToM: Multi-modal Multi-Agent Theory of MindHaojun Shi, Suyu Ye, Xinyu Fang, Chuanyang Jin 等AAAI 2025 · 被引用 48 次
- SimpleToM: Exposing the Gap between Explicit ToM Inference and Implicit ToM Application in LLMsYuling Gu, Oyvind Tafjord, Hyunwoo Kim, Jared Moore 等ICLR 2026 · 被引用 39 次
- AutoToM: Scaling Model-based Mental Inference via Automated Agent ModelingZhining Zhang, Chuanyang Jin, Mung Yao Jia, Shunchi Zhang 等NeurIPS 2025 · 被引用 30 次
- Privacy Reasoning in Ambiguous ContextsRen Yi, Octavian Suciu, Adrià Gascón, Sarah Meiklejohn 等NeurIPS 2025 · 被引用 15 次
它引用的顶会 Paper10
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- UL2: Unifying Language Learning ParadigmsYi Tay, Mostafa Dehghani, Vinh Q. Tran, Xavier Garcia 等ICLR 2023 · 被引用 97 次
- Neural Theory-of-Mind? On the Limits of Social Intelligence in Large LMsMaarten Sap, Ronan Le Bras, Daniel Fried, Yejin ChoiEMNLP 2022 · 被引用 92 次
相关 Paper
- Perceptions to Beliefs: Exploring Precursory Inferences for Theory of Mind in Large Language ModelsChani Jung, Dongkwan Kim, Jiho Jin, Jiseon Kim 等EMNLP 2024 · 被引用 2 次
- ToMATO: Verbalizing the Mental States of Role-Playing LLMs for Benchmarking Theory of MindKazutoshi Shinoda, Nobukatsu Hojo, Kyosuke Nishida, Saki Mizuno 等AAAI 2025 · 被引用 10 次
- CogToM: A Comprehensive Theory of Mind Benchmark inspired by Human Cognition for Large Language ModelsHaibo Tong, Zeyang Yue, Feifei Zhao, Erliang Lin 等ACL 2026
- RecToM: A Benchmark for Evaluating Machine Theory of Mind in LLM-based Conversational Recommender SystemsMengfan Li, Xuanhua Shi, Yang DengAAAI 2026
- Theory of Mind in Large Language Models: Assessment and EnhancementRuirui Chen, Weifeng Jiang, Chengwei Qin, Cheston TanACL 2025
