Make LLMs See Like Investigators, Not Just Think More: The Role of Structured Analysis in Investigative Reasoning
Jaewook Lee, Myeong-Cheol Kang, Jong-hun Shin
摘要
Criminal investigators and intelligence analysts have developed structured analytic techniques to evaluate competing hypotheses under incomplete information. This study examines whether such human expert investigative methodologies are also effective for narrativebased culprit inference in large language models (LLMs). Focusing on the task of analyzing evidence from complex narratives and identifying the perpetrator among suspects, we conducted experiments on 10 LLMs using the MuSR murder mystery benchmark. The PRISM framework, which applies investigative techniques, consistently outperformed existing general-purpose strategies across all models, with its effectiveness manifesting regardless of model scale. Ablation studies revealed that the hypothesis structuring stage is particularly crucial, accounting for 89% of the methodological improvement beyond information filtering. This suggests that domain-specific structures that specify "what to analyze" are more effective in LLM reasoning than simply increasing the number of reasoning paths.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model SocietyGuohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin 等NeurIPS 2023 · 被引用 1,975 次
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales 等ICML 2023 · 被引用 970 次
- Analyzing Key Factors Influencing Emotion Prediction Performance of VLLMs in Conversational ContextsJaewook Lee, Yeajin Jang, Hongjin Kim, Woojin Lee 等EMNLP 2024 · 被引用 4 次
- Small Changes, Big Impact: How Manipulating a Few Neurons Can Drastically Alter LLM AggressionJaewook Lee, Junseo Jang, Oh-Woog Kwon, Harksoo KimACL 2025
相关 Paper
- MuSR: Testing the Limits of Chain-of-thought with Multistep Soft ReasoningZayne Sprague, Xi Ye, Kaj Bostrom, Swarat Chaudhuri 等ICLR 2024 · 被引用 172 次
- Creativity or Brute Force? Using Brainteasers as a Window into the Problem-Solving Abilities of Large Language ModelsSophia Simeng Han, Howard Dai, Stephen Xia, Grant Zhang 等NeurIPS 2025 · 被引用 2 次
- How LLMs Comprehend Temporal Meaning in Narratives: A Case Study in Cognitive Evaluation of LLMsKarin de Langis, Jong Inn Park, Andreas Schramm, Bin Hu 等ACL 2025
- Prism: A Framework for Decoupling and Assessing the Capabilities of VLMsYuxuan Qiao, Haodong Duan, Xinyu Fang, Junming Yang 等NeurIPS 2024 · 被引用 49 次
- TurnaboutLLM: A Deductive Reasoning Benchmark from Detective GamesYuan Yuan, Muyu He, Muhammad Adil Shahid, Ziyang Li 等EMNLP 2025
