InterIntent: Investigating Social Intelligence of LLMs via Intention Understanding in an Interactive Game Context
Ziyi Liu, Abhishek Anand, Pei Zhou, Jen-tse Huang, Jieyu Zhao
Abstract
Large language models (LLMs) have demonstrated the potential to mimic human social intelligence. However, most studies focus on simplistic and static self-report or performancebased tests, which limits the depth and validity of the analysis. In this paper, we developed a novel framework, INTERINTENT, to assess LLMs' social intelligence by mapping their ability to understand and manage intentions in a game setting. We focus on four dimensions of social intelligence: situational awareness, selfregulation, self-awareness, and theory of mind. Each dimension is linked to a specific game task: intention selection, intention following, intention summarization, and intention guessing. Our findings indicate that while LLMs exhibit high proficiency in selecting intentions, achieving an accuracy of 88%, their ability to infer the intentions of others is significantly weaker, trailing human performance by 20%. Additionally, game performance correlates with intention understanding, highlighting the importance of the four components towards success in this game. These findings underline the crucial role of intention understanding in evaluating LLMs' social intelligence and highlight the potential of using social deduction games as a complex testbed to enhance LLM evaluation. INTERINTENT contributes a structured approach to bridging the evaluation gap in social intelligence within multiplayer games. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6065e0bc-c11c-45a2-acfa-3cdc8c2b5a0aCited by top-tier papers6
- SimpleToM: Exposing the Gap between Explicit ToM Inference and Implicit ToM Application in LLMsYuling Gu, Oyvind Tafjord, Hyunwoo Kim, Jared Moore et al.ICLR 2026 · 39 citations
- Bayesian Social Deduction with Graph-Informed Language ModelsShahab Rahimirad, Guven Gergerli, Lucia Romero, Angela Qian et al.ACL 2026 · 4 citations
- PRISON: Unmasking the Criminal Potential of Large Language ModelsXinyi Wu, Geng Hong, Pei Chen, Yueyue Chen et al.ICLR 2026 · 3 citations
- NewsInterview: a Dataset and a Playground to Evaluate LLMs' Grounding Gap via Informational InterviewsAlexander Spangher, Michael Lu, Sriya Kalyan, Hyundong Justin Cho et al.ACL 2025 · 2 citations
- InMind: Evaluating LLMs in Capturing and Applying Individual Human Reasoning StylesZizhen Li, Chuanhao Li, Yibin Wang, Qi Chen et al.EMNLP 2025
Builds on9
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- Language Agents with Reinforcement Learning for Strategic Play in the Werewolf GameZelai Xu, Chao Yu, Fei Fang, Yu Wang et al.ICML 2024 · 145 citations
- Neural Theory-of-Mind? On the Limits of Social Intelligence in Large LMsMaarten Sap, Ronan Le Bras, Daniel Fried, Yejin ChoiEMNLP 2022 · 92 citations
- On the Humanity of Conversational AI: Evaluating the Psychological Portrayal of LLMsJen-tse Huang, Wenxuan Wang, Eric John Li, Man Ho Lam et al.ICLR 2024 · 85 citations
- Minding Language Models' (Lack of) Theory of Mind: A Plug-and-Play Multi-Character Belief TrackerMelanie Sclar, Sachin Kumar, Peter West, Alane Suhr et al.ACL 2023 · 21 citations
Related papers
- MAgIC: Investigation of Large Language Model Powered Multi-Agent in Cognition, Adaptability, Rationality and CollaborationLin Xu, Zhiyuan Hu, Daquan Zhou, Hongyu Ren et al.EMNLP 2024 · 7 citations
- clembench: Using Game Play to Evaluate Chat-Optimized Language Models as Conversational AgentsKranti Chalamalasetti, Jana Götze, Sherzod Hakimov, Brielen Madureira et al.EMNLP 2023 · 6 citations
- ToMBench: Benchmarking Theory of Mind in Large Language ModelsZhuang Chen, Jincenzi Wu, Jinfeng Zhou, Bosi Wen et al.ACL 2024 · 6 citations
- MetaMind: Modeling Human Social Thoughts with Metacognitive Multi-Agent SystemsXuanming Zhang, Yuxuan Chen, Samuel (Min-Hsuan) Yeh, Sharon LiNeurIPS 2025 · 14 citations
- Theory of Mind in Large Language Models: Assessment and EnhancementRuirui Chen, Weifeng Jiang, Chengwei Qin, Cheston TanACL 2025
