RFS-Guard: Detecting Reasoning Hallucinations via Cross-Phase Routing Focus in Large Reasoning Models
Zihang Liu, Zhouhua Fang, Hui Liu, Zhiwei Liu, Yong Li, Haishuai Wang
摘要
Large reasoning models (LRMs) achieve strong performance on complex tasks by generating intermediate reasoning before the final answer, yet they remain prone to reasoning hallucinations such as subtle arithmetic or constraintviolation errors. Prior hallucination detectors often rely on external verification or local tokenlevel signals, which are limited for LRMs and largely overlook whether the cross-phase information flow from reasoning to answering is structurally robust. We propose Routing Focus Score (RFS), a step-level indicator that measures how strongly cross-step attention routing aligns with semantic proximity derived from hidden-state cosine similarity. We further design RFS-Guard, a lightweight hallucination detection framework based on RFS. Empirically, we observe that higher reasoning-answer RFS is consistently associated with higher hallucination risk, suggesting a routing-collapse failure mode where models might prefer selfconfirmation loops and suppress the ability to audit their own generations. Experimental results across multiple domains and models demonstrate the superiority of RFS-Guard for detecting and localizing hallucinations in LRMs without requiring external tools or repeated sampling.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister 等NeurIPS 2023 · 被引用 1,549 次
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language ModelsPotsawee Manakul, Adian Liusie, Mark J. F. GalesEMNLP 2023 · 被引用 331 次
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text GenerationSewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis 等EMNLP 2023 · 被引用 225 次
- SelfCheck: Using LLMs to Zero-Shot Check Their Own Step-by-Step ReasoningNing Miao, Yee Whye Teh, Tom RainforthICLR 2024 · 被引用 195 次
- Long-form factuality in large language modelsJerry Wei, Chengrun Yang, Xinying Song, Yifeng Lu 等NeurIPS 2024 · 被引用 182 次
相关 Paper
- Mechanistic Detection and Mitigation of Hallucination in Large Reasoning ModelsZhongxiang Sun, Qipeng Wang, Haoyu Wang, Xiao Zhang 等ICLR 2026 · 被引用 30 次
- Joint Evaluation of Answer and Reasoning Consistency for Hallucination Detection in Large Reasoning ModelsChangyue Wang, Weihang Su, Qingyao Ai, Yiqun LiuAAAI 2026 · 被引用 13 次
- Hallucination Detection from Structural Reasoning ModelJianbo Sun, Pengkun YangICML 2026
- Harnessing Reasoning Trajectories for Hallucination Detection via Answer-agreement Representation ShapingJianxiong Zhang, Bing Guo, Yuming Jiang, Haobo Wang 等ICML 2026 · 被引用 2 次
- From Reasoning to Answer: Empirical, Attention-Based and Mechanistic Insights into Distilled DeepSeek R1 ModelsJue Zhang, Qingwei Lin, Saravan Rajmohan, Dongmei ZhangEMNLP 2025
