MuMA-ToM: Multi-modal Multi-Agent Theory of Mind
Haojun Shi, Suyu Ye, Xinyu Fang, Chuanyang Jin, Leyla Isik, Yen-Ling Kuo, Tianmin Shu
Abstract
Understanding people's social interactions in complex real-world scenarios often relies on intricate mental reasoning. To truly understand how and why people interact with one another, we must infer the underlying mental states that give rise to the social interactions, i.e., Theory of Mind reasoning in multi-agent interactions. Additionally, social interactions are often multi-modal -- we can watch people's actions, hear their conversations, and/or read about their past behaviors. For AI systems to successfully and safely interact with people in real-world environments, they also need to understand people's mental states as well as their inferences about each other's mental states based on multi-modal information about their interactions. For this, we introduce MuMA-ToM, a Multi-modal Multi-Agent Theory of Mind benchmark. MuMA-ToM is the first multi-modal Theory of Mind benchmark that evaluates mental reasoning in embodied multi-agent interactions. In MuMA-ToM, we provide video and text descriptions of people's multi-modal behavior in realistic household environments. Based on the context, we then ask questions about people's goals, beliefs, and beliefs about others' goals. We validated MuMA-ToM in a human experiment and provided a human baseline. We also proposed a novel multi-modal, multi-agent ToM model, LIMP (Language model-based Inverse Multi-agent Planning). Our experimental results show that LIMP significantly outperforms state-of-the-art methods, including large multi-modal models (e.g., GPT-4o, Gemini-1.5 Pro) and a recent multi-modal ToM model, BIP-ALM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 310cb4e9-160b-411b-b23a-ad1cf4eaf9fdCited by top-tier papers18
- AutoToM: Scaling Model-based Mental Inference via Automated Agent ModelingZhining Zhang, Chuanyang Jin, Mung Yao Jia, Shunchi Zhang et al.NeurIPS 2025 · 30 citations
- Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human StatesYang Xiao, Jiashuo Wang, Qiancheng Xu, Changhe Song et al.ACL 2025 · 12 citations
- UniM: A Unified Any-to-Any Interleaved Multimodal BenchmarkYanlin Li, Minghui Guo, Kaiwen Zhang, Shize Zhang et al.CVPR 2026 · 10 citations
- MindPower: Enabling Theory-of-Mind Reasoning in VLM-based Embodied AgentsRuoxuan Zhang, Qiyun Zheng, Zhiyu Zhou, Ziqi Liao et al.CVPR 2026 · 6 citations
- Hierarchical Attacks for Multi-Modal Multi-Agent ReasoningHao Zhou, Tiru Wu, Yan Jiang, Wanqi Zhou et al.CVPR 2026 · 1 citation
Builds on13
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Watch-And-Help: A Challenge for Social Perception and Human-AI CollaborationXavier Puig, Tianmin Shu, Shuang Li, Zilin Wang et al.ICLR 2021 · 170 citations
- AGENT: A Benchmark for Core Psychological ReasoningTianmin Shu, Abhishek Bhandwaldar, Chuang Gan, Kevin A. Smith et al.ICML 2021 · 79 citations
- Baby Intuitions Benchmark (BIB): Discerning the goals, preferences, and actions of othersKanishk Gandhi, Gala Stojnic, Brenden M. Lake, Moira R. DillonNeurIPS 2021 · 59 citations
- PHASE: PHysically-grounded Abstract Social Events for Machine Social PerceptionAviv Netanyahu, Tianmin Shu, Boris Katz, Andrei Barbu et al.AAAI 2021 · 44 citations
Related papers
- MMToM-QA: Multimodal Theory of Mind Question AnsweringChuanyang Jin, Yutong Wu, Jing Cao, Jiannan Xiang et al.ACL 2024 · 8 citations
- MetaMind: Modeling Human Social Thoughts with Metacognitive Multi-Agent SystemsXuanming Zhang, Yuxuan Chen, Samuel (Min-Hsuan) Yeh, Sharon LiNeurIPS 2025 · 14 citations
- GroupToM-Bench: Benchmarking Group Theory of Mind and Nonlinear Social Emergence in MLLMsWeidong Tang, Jierui Li, Yueling Hou, Zihan Mei et al.ACL 2026
- Tracing Belief-Driven Thoughts with Theory-of-Mind Agents: An Opinion Analysis FrameworkJintao Wen, Yunfeng Ning, Hankun Kang, Xin Miao et al.WWW 2026
- Overcoming Multi-step Complexity in Multimodal Theory-of-Mind Reasoning: A Scalable Bayesian PlannerChunhui Zhang, Zhongyu Ouyang, Kwonjoon Lee, Nakul Agarwal et al.ICML 2025
