ToM-SSI: Evaluating Theory of Mind in Situated Social Interactions
Matteo Bortoletto, Constantin Ruhdorfer, Andreas Bulling
摘要
Most existing Theory of Mind (ToM) benchmarks for foundation models rely on variations of the Sally-Anne test, offering only a very limited perspective on ToM and neglecting the complexity of human social interactions. To address this gap, we propose ToM-SSI: a new benchmark specifically designed to test ToM capabilities in environments rich with social interactions and spatial dynamics. While current ToM benchmarks are limited to text-only or dyadic interactions, ToM-SSI is multimodal and includes group interactions of up to four agents that communicate and move in situated environments. This unique design allows us to study, for the first time, mixed cooperativeobstructive settings and reasoning about multiple agents' mental state in parallel, thus capturing a wider range of social cognition than existing benchmarks. Our evaluations reveal that the current models' performance is still severely limited, especially in these new tasks, highlighting critical gaps for future research. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper16
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Baby Intuitions Benchmark (BIB): Discerning the goals, preferences, and actions of othersKanishk Gandhi, Gala Stojnic, Brenden M. Lake, Moira R. DillonNeurIPS 2021 · 被引用 59 次
- MuMA-ToM: Multi-modal Multi-Agent Theory of MindHaojun Shi, Suyu Ye, Xinyu Fang, Chuanyang Jin 等AAAI 2025 · 被引用 48 次
- SimpleToM: Exposing the Gap between Explicit ToM Inference and Implicit ToM Application in LLMsYuling Gu, Oyvind Tafjord, Hyunwoo Kim, Jared Moore 等ICLR 2026 · 被引用 39 次
- MindCraft: Theory of Mind Modeling for Situated Dialogue in Collaborative TasksCristian-Paul Bara, Sky CH-Wang, Joyce ChaiEMNLP 2021 · 被引用 28 次
相关 Paper
- FANToM: A Benchmark for Stress-testing Machine Theory of Mind in InteractionsHyunwoo Kim, Melanie Sclar, Xuhui Zhou, Ronan Le Bras 等EMNLP 2023 · 被引用 21 次
- Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human StatesYang Xiao, Jiashuo Wang, Qiancheng Xu, Changhe Song 等ACL 2025 · 被引用 12 次
- RecToM: A Benchmark for Evaluating Machine Theory of Mind in LLM-based Conversational Recommender SystemsMengfan Li, Xuanhua Shi, Yang DengAAAI 2026
- ToMBench: Benchmarking Theory of Mind in Large Language ModelsZhuang Chen, Jincenzi Wu, Jinfeng Zhou, Bosi Wen 等ACL 2024 · 被引用 6 次
- OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language ModelsHainiu Xu, Runcong Zhao, Lixing Zhu, Jinhua Du 等ACL 2024
