InterIntent: Investigating Social Intelligence of LLMs via Intention Understanding in an Interactive Game Context
Ziyi Liu, Abhishek Anand, Pei Zhou, Jen-tse Huang, Jieyu Zhao
摘要
Large language models (LLMs) have demonstrated the potential to mimic human social intelligence. However, most studies focus on simplistic and static self-report or performancebased tests, which limits the depth and validity of the analysis. In this paper, we developed a novel framework, INTERINTENT, to assess LLMs' social intelligence by mapping their ability to understand and manage intentions in a game setting. We focus on four dimensions of social intelligence: situational awareness, selfregulation, self-awareness, and theory of mind. Each dimension is linked to a specific game task: intention selection, intention following, intention summarization, and intention guessing. Our findings indicate that while LLMs exhibit high proficiency in selecting intentions, achieving an accuracy of 88%, their ability to infer the intentions of others is significantly weaker, trailing human performance by 20%. Additionally, game performance correlates with intention understanding, highlighting the importance of the four components towards success in this game. These findings underline the crucial role of intention understanding in evaluating LLMs' social intelligence and highlight the potential of using social deduction games as a complex testbed to enhance LLM evaluation. INTERINTENT contributes a structured approach to bridging the evaluation gap in social intelligence within multiplayer games. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- SimpleToM: Exposing the Gap between Explicit ToM Inference and Implicit ToM Application in LLMsYuling Gu, Oyvind Tafjord, Hyunwoo Kim, Jared Moore 等ICLR 2026 · 被引用 39 次
- Bayesian Social Deduction with Graph-Informed Language ModelsShahab Rahimirad, Guven Gergerli, Lucia Romero, Angela Qian 等ACL 2026 · 被引用 4 次
- PRISON: Unmasking the Criminal Potential of Large Language ModelsXinyi Wu, Geng Hong, Pei Chen, Yueyue Chen 等ICLR 2026 · 被引用 3 次
- NewsInterview: a Dataset and a Playground to Evaluate LLMs' Grounding Gap via Informational InterviewsAlexander Spangher, Michael Lu, Sriya Kalyan, Hyundong Justin Cho 等ACL 2025 · 被引用 2 次
- InMind: Evaluating LLMs in Capturing and Applying Individual Human Reasoning StylesZizhen Li, Chuanhao Li, Yibin Wang, Qi Chen 等EMNLP 2025
它引用的顶会 Paper9
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris 等UIST 2023 · 被引用 1,882 次
- Language Agents with Reinforcement Learning for Strategic Play in the Werewolf GameZelai Xu, Chao Yu, Fei Fang, Yu Wang 等ICML 2024 · 被引用 145 次
- Neural Theory-of-Mind? On the Limits of Social Intelligence in Large LMsMaarten Sap, Ronan Le Bras, Daniel Fried, Yejin ChoiEMNLP 2022 · 被引用 92 次
- On the Humanity of Conversational AI: Evaluating the Psychological Portrayal of LLMsJen-tse Huang, Wenxuan Wang, Eric John Li, Man Ho Lam 等ICLR 2024 · 被引用 85 次
- Minding Language Models' (Lack of) Theory of Mind: A Plug-and-Play Multi-Character Belief TrackerMelanie Sclar, Sachin Kumar, Peter West, Alane Suhr 等ACL 2023 · 被引用 21 次
相关 Paper
- MAgIC: Investigation of Large Language Model Powered Multi-Agent in Cognition, Adaptability, Rationality and CollaborationLin Xu, Zhiyuan Hu, Daquan Zhou, Hongyu Ren 等EMNLP 2024 · 被引用 7 次
- clembench: Using Game Play to Evaluate Chat-Optimized Language Models as Conversational AgentsKranti Chalamalasetti, Jana Götze, Sherzod Hakimov, Brielen Madureira 等EMNLP 2023 · 被引用 6 次
- ToMBench: Benchmarking Theory of Mind in Large Language ModelsZhuang Chen, Jincenzi Wu, Jinfeng Zhou, Bosi Wen 等ACL 2024 · 被引用 6 次
- MetaMind: Modeling Human Social Thoughts with Metacognitive Multi-Agent SystemsXuanming Zhang, Yuxuan Chen, Samuel (Min-Hsuan) Yeh, Sharon LiNeurIPS 2025 · 被引用 14 次
- Theory of Mind in Large Language Models: Assessment and EnhancementRuirui Chen, Weifeng Jiang, Chengwei Qin, Cheston TanACL 2025
