Hallucination Detection from Structural Reasoning Model
Jianbo Sun, Pengkun Yang
摘要
Hallucinations pose a key challenge for large language models, and chain-of-thought prompting exposes intermediate reasoning but usually treats traces as linear sequences, making crossstep dependencies and unsupported intermediate claims difficult to identify. We propose a structural reasoning model to describe interactions among local reasoning steps. To detect hallucinations, we extract a directed acyclic reasoning graph over conditions and intermediate claims, verify each claim against its parent nodes, and aggregate the step signals with a simple mass-flow rule. Under a probabilistic erasure-gate abstraction, we interpret this aggregation as measuring information loss along the reasoning graph. Experiments on GSM8K, MATH, HumanEval, and HotpotQA show that the proposed method is most advantageous on longer, dependency-rich reasoning traces such as math and code generation, while remaining competitive on shorter factual QA; these results provide a structured perspective on chain-of-thought evaluation. All code and data are available at https://github.com/ soncheinbok/FlowScore.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 被引用 1,792 次
- Graph of Thoughts: Solving Elaborate Problems with Large Language ModelsMaciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger 等AAAI 2024 · 被引用 1,292 次
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang 等EMNLP 2023 · 被引用 549 次
相关 Paper
- RFS-Guard: Detecting Reasoning Hallucinations via Cross-Phase Routing Focus in Large Reasoning ModelsZihang Liu, Zhouhua Fang, Hui Liu, Zhiwei Liu 等ACL 2026
- Mind the Gap: Catching Hallucinations via Evidence Drop on the Reasoning ManifoldQunJie Chen, Yufei Chen, Xiaodong Yue, Linye LiICML 2026
- Mapping the Minds of LLMs: A Graph-Based Analysis of Reasoning LLMsZhen Xiong, Yujun Cai, Zhecheng Li, Yiwei WangEMNLP 2025
- Chain-of-Thought as a Lens: Evaluating Structured Reasoning Alignment between Human Preferences and Large Language ModelsBoxuan Wang, Zhuoyun Li, Xinmiao Huang, Xiaowei Huang 等ACL 2026 · 被引用 2 次
- Understanding Chain-of-Thought in LLMs through Information TheoryJean-Francois Ton, Muhammad Faaiz Taufiq, Yang LiuICML 2025
