Hallucination Detection from Structural Reasoning Model
Jianbo Sun, Pengkun Yang
Abstract
Hallucinations pose a key challenge for large language models, and chain-of-thought prompting exposes intermediate reasoning but usually treats traces as linear sequences, making crossstep dependencies and unsupported intermediate claims difficult to identify. We propose a structural reasoning model to describe interactions among local reasoning steps. To detect hallucinations, we extract a directed acyclic reasoning graph over conditions and intermediate claims, verify each claim against its parent nodes, and aggregate the step signals with a simple mass-flow rule. Under a probabilistic erasure-gate abstraction, we interpret this aggregation as measuring information loss along the reasoning graph. Experiments on GSM8K, MATH, HumanEval, and HotpotQA show that the proposed method is most advantageous on longer, dependency-rich reasoning traces such as math and code generation, while remaining competitive on shorter factual QA; these results provide a structured perspective on chain-of-thought evaluation. All code and data are available at https://github.com/ soncheinbok/FlowScore.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4cac72e4-406a-416a-938d-15a6d7e70710Builds on11
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 1,792 citations
- Graph of Thoughts: Solving Elaborate Problems with Large Language ModelsMaciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger et al.AAAI 2024 · 1,292 citations
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang et al.EMNLP 2023 · 549 citations
Related papers
- RFS-Guard: Detecting Reasoning Hallucinations via Cross-Phase Routing Focus in Large Reasoning ModelsZihang Liu, Zhouhua Fang, Hui Liu, Zhiwei Liu et al.ACL 2026
- Mind the Gap: Catching Hallucinations via Evidence Drop on the Reasoning ManifoldQunJie Chen, Yufei Chen, Xiaodong Yue, Linye LiICML 2026
- Mapping the Minds of LLMs: A Graph-Based Analysis of Reasoning LLMsZhen Xiong, Yujun Cai, Zhecheng Li, Yiwei WangEMNLP 2025
- Chain-of-Thought as a Lens: Evaluating Structured Reasoning Alignment between Human Preferences and Large Language ModelsBoxuan Wang, Zhuoyun Li, Xinmiao Huang, Xiaowei Huang et al.ACL 2026 · 2 citations
- Understanding Chain-of-Thought in LLMs through Information TheoryJean-Francois Ton, Muhammad Faaiz Taufiq, Yang LiuICML 2025
