InMind: Evaluating LLMs in Capturing and Applying Individual Human Reasoning Styles
Zizhen Li, Chuanhao Li, Yibin Wang, Qi Chen, Diping Song, Yukang Feng, Jianwen Sun, Jiaxin Ai, Fanrui Zhang, Mingzhu Sun, Kaipeng Zhang
摘要
LLMs have shown strong performance on human-centric reasoning tasks. While previous evaluations have explored whether LLMs can infer intentions or detect deception, they often overlook the individualized reasoning styles that influence how people interpret and act in social contexts. Social deduction games (SDGs) provide a natural testbed for evaluating individualized reasoning styles, where different players may adopt diverse but contextually valid reasoning strategies under identical conditions. To address this, we introduce InMind, a cognitively grounded evaluation framework designed to assess whether LLMs can capture and apply personalized reasoning styles in SDGs. InMind enhances structured gameplay data with round-level strategy traces and post-game reflections, collected under both Observer and Participant modes. It supports four cognitively motivated tasks that jointly evaluate both static alignment and dynamic adaptation. As a case study, we apply InMind to the game Avalon, evaluating 11 state-of-the-art LLMs. 1 Generalpurpose LLMs, even GPT-4o frequently rely on lexical cues, struggling to anchor reflections in temporal gameplay or adapt to evolving strategies. In contrast, reasoning-enhanced LLMs like DeepSeek-R1 exhibit early signs of stylesensitive reasoning. These findings reveal key limitations in current LLMs' capacity for individualized, adaptive reasoning, and position InMind as a step toward cognitively aligned human-AI interaction. Emerging research (Strachan et al., 2024; Mittelstädt et al., 2024) further highlights their promising performance in human-centric tasks, including social commonsense inference, intention recognition, and belief attribution. Beyond these capabilities, recent studies suggest that LLMs may exhibit early signs of Theory of Mind (ToM)-the ability to represent and reason about others ' beliefs, desires, and intentions (Sarıtaş et al., 2025; Kim et al., 2025) . Understanding and evaluating such high-level cognitive traits is critical for advancing LLMs toward artificial general intelligence (AGI), and potentially, artificial superintelligence. Existing benchmarks attempt to assess ToMlike reasoning through tasks such as intent classification (Liu et al., 2024) , false-belief attribu-
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- StructBERT: Incorporating Language Structures into Pre-training for Deep Language UnderstandingWei Wang, Bin Bi, Ming Yan, Chen Wu 等ICLR 2020 · 被引用 297 次
- InterIntent: Investigating Social Intelligence of LLMs via Intention Understanding in an Interactive Game ContextZiyi Liu, Abhishek Anand, Pei Zhou, Jen-tse Huang 等EMNLP 2024 · 被引用 2 次
- MultiMind: Enhancing Werewolf Agents with Multimodal Reasoning and Theory of MindZheng Zhang, Nuoqian Xiao, Qi Chai, Deheng Ye 等ACM MM 2025 · 被引用 1 次
- OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language ModelsHainiu Xu, Runcong Zhao, Lixing Zhu, Jinhua Du 等ACL 2024
相关 Paper
- Bayesian Social Deduction with Graph-Informed Language ModelsShahab Rahimirad, Guven Gergerli, Lucia Romero, Angela Qian 等ACL 2026 · 被引用 4 次
- The Decrypto Benchmark for Multi-Agent Reasoning and Theory of MindAndrei Lupu, Timon Willi, Jakob FoersterICML 2026 · 被引用 2 次
- ToMBench: Benchmarking Theory of Mind in Large Language ModelsZhuang Chen, Jincenzi Wu, Jinfeng Zhou, Bosi Wen 等ACL 2024 · 被引用 6 次
- LLM Strategic Reasoning: Agentic Study through Behavioral Game TheoryJingru Jia, Zehua Yuan, Junhao Pan, Paul McNamara 等NeurIPS 2025 · 被引用 23 次
- MetaMind: Modeling Human Social Thoughts with Metacognitive Multi-Agent SystemsXuanming Zhang, Yuxuan Chen, Samuel (Min-Hsuan) Yeh, Sharon LiNeurIPS 2025 · 被引用 14 次
