InMind: Evaluating LLMs in Capturing and Applying Individual Human Reasoning Styles
Zizhen Li, Chuanhao Li, Yibin Wang, Qi Chen, Diping Song, Yukang Feng, Jianwen Sun, Jiaxin Ai, Fanrui Zhang, Mingzhu Sun, Kaipeng Zhang
Abstract
LLMs have shown strong performance on human-centric reasoning tasks. While previous evaluations have explored whether LLMs can infer intentions or detect deception, they often overlook the individualized reasoning styles that influence how people interpret and act in social contexts. Social deduction games (SDGs) provide a natural testbed for evaluating individualized reasoning styles, where different players may adopt diverse but contextually valid reasoning strategies under identical conditions. To address this, we introduce InMind, a cognitively grounded evaluation framework designed to assess whether LLMs can capture and apply personalized reasoning styles in SDGs. InMind enhances structured gameplay data with round-level strategy traces and post-game reflections, collected under both Observer and Participant modes. It supports four cognitively motivated tasks that jointly evaluate both static alignment and dynamic adaptation. As a case study, we apply InMind to the game Avalon, evaluating 11 state-of-the-art LLMs. 1 Generalpurpose LLMs, even GPT-4o frequently rely on lexical cues, struggling to anchor reflections in temporal gameplay or adapt to evolving strategies. In contrast, reasoning-enhanced LLMs like DeepSeek-R1 exhibit early signs of stylesensitive reasoning. These findings reveal key limitations in current LLMs' capacity for individualized, adaptive reasoning, and position InMind as a step toward cognitively aligned human-AI interaction. Emerging research (Strachan et al., 2024; Mittelstädt et al., 2024) further highlights their promising performance in human-centric tasks, including social commonsense inference, intention recognition, and belief attribution. Beyond these capabilities, recent studies suggest that LLMs may exhibit early signs of Theory of Mind (ToM)-the ability to represent and reason about others ' beliefs, desires, and intentions (Sarıtaş et al., 2025; Kim et al., 2025) . Understanding and evaluating such high-level cognitive traits is critical for advancing LLMs toward artificial general intelligence (AGI), and potentially, artificial superintelligence. Existing benchmarks attempt to assess ToMlike reasoning through tasks such as intent classification (Liu et al., 2024) , false-belief attribu-
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 97b273c6-a7ac-4f8e-8056-a5fb7e4e9d06Builds on4
- StructBERT: Incorporating Language Structures into Pre-training for Deep Language UnderstandingWei Wang, Bin Bi, Ming Yan, Chen Wu et al.ICLR 2020 · 297 citations
- InterIntent: Investigating Social Intelligence of LLMs via Intention Understanding in an Interactive Game ContextZiyi Liu, Abhishek Anand, Pei Zhou, Jen-tse Huang et al.EMNLP 2024 · 2 citations
- MultiMind: Enhancing Werewolf Agents with Multimodal Reasoning and Theory of MindZheng Zhang, Nuoqian Xiao, Qi Chai, Deheng Ye et al.ACM MM 2025 · 1 citation
- OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language ModelsHainiu Xu, Runcong Zhao, Lixing Zhu, Jinhua Du et al.ACL 2024
Related papers
- Bayesian Social Deduction with Graph-Informed Language ModelsShahab Rahimirad, Guven Gergerli, Lucia Romero, Angela Qian et al.ACL 2026 · 4 citations
- The Decrypto Benchmark for Multi-Agent Reasoning and Theory of MindAndrei Lupu, Timon Willi, Jakob FoersterICML 2026 · 2 citations
- ToMBench: Benchmarking Theory of Mind in Large Language ModelsZhuang Chen, Jincenzi Wu, Jinfeng Zhou, Bosi Wen et al.ACL 2024 · 6 citations
- LLM Strategic Reasoning: Agentic Study through Behavioral Game TheoryJingru Jia, Zehua Yuan, Junhao Pan, Paul McNamara et al.NeurIPS 2025 · 23 citations
- MetaMind: Modeling Human Social Thoughts with Metacognitive Multi-Agent SystemsXuanming Zhang, Yuxuan Chen, Samuel (Min-Hsuan) Yeh, Sharon LiNeurIPS 2025 · 14 citations
