Make LLMs See Like Investigators, Not Just Think More: The Role of Structured Analysis in Investigative Reasoning
Jaewook Lee, Myeong-Cheol Kang, Jong-hun Shin
Abstract
Criminal investigators and intelligence analysts have developed structured analytic techniques to evaluate competing hypotheses under incomplete information. This study examines whether such human expert investigative methodologies are also effective for narrativebased culprit inference in large language models (LLMs). Focusing on the task of analyzing evidence from complex narratives and identifying the perpetrator among suspects, we conducted experiments on 10 LLMs using the MuSR murder mystery benchmark. The PRISM framework, which applies investigative techniques, consistently outperformed existing general-purpose strategies across all models, with its effectiveness manifesting regardless of model scale. Ablation studies revealed that the hypothesis structuring stage is particularly crucial, accounting for 89% of the methodological improvement beyond information filtering. This suggests that domain-specific structures that specify "what to analyze" are more effective in LLM reasoning than simply increasing the number of reasoning paths.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on7
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model SocietyGuohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin et al.NeurIPS 2023 · 1,975 citations
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales et al.ICML 2023 · 970 citations
- Analyzing Key Factors Influencing Emotion Prediction Performance of VLLMs in Conversational ContextsJaewook Lee, Yeajin Jang, Hongjin Kim, Woojin Lee et al.EMNLP 2024 · 4 citations
- Small Changes, Big Impact: How Manipulating a Few Neurons Can Drastically Alter LLM AggressionJaewook Lee, Junseo Jang, Oh-Woog Kwon, Harksoo KimACL 2025
Related papers
- MuSR: Testing the Limits of Chain-of-thought with Multistep Soft ReasoningZayne Sprague, Xi Ye, Kaj Bostrom, Swarat Chaudhuri et al.ICLR 2024 · 172 citations
- Creativity or Brute Force? Using Brainteasers as a Window into the Problem-Solving Abilities of Large Language ModelsSophia Simeng Han, Howard Dai, Stephen Xia, Grant Zhang et al.NeurIPS 2025 · 2 citations
- How LLMs Comprehend Temporal Meaning in Narratives: A Case Study in Cognitive Evaluation of LLMsKarin de Langis, Jong Inn Park, Andreas Schramm, Bin Hu et al.ACL 2025
- Prism: A Framework for Decoupling and Assessing the Capabilities of VLMsYuxuan Qiao, Haodong Duan, Xinyu Fang, Junming Yang et al.NeurIPS 2024 · 49 citations
- TurnaboutLLM: A Deductive Reasoning Benchmark from Detective GamesYuan Yuan, Muyu He, Muhammad Adil Shahid, Ziyang Li et al.EMNLP 2025
