Harnessing Reasoning Trajectories for Hallucination Detection via Answer-agreement Representation Shaping
Jianxiong Zhang, Bing Guo, Yuming Jiang, Haobo Wang, Bo An, Sean Du
Abstract
Large reasoning models (LRMs) often generate long, seemingly coherent reasoning traces yet still produce incorrect answers, making hallucination detection challenging. Although trajectories contain useful signals, directly using trace text or vanilla hidden states for detection is brittle: traces vary in form and detectors can overfit to superficial patterns rather than answer validity. We introduce Answer-agreement Representation Shaping (ARS), which learns detection-friendly trace-conditioned representations by explicitly encoding answer stability. ARS generates counterfactual answers through small latent interventions, specifically, perturbing the trace-boundary embedding, and labels each perturbation by whether the resulting answer agrees with the original. It then learns representations that bring answer-agreeing states together and separate answer-disagreeing ones, exposing latent instability indicative of hallucination risk. The shaped embeddings are plug-and-play with existing embedding-based detectors and require no human annotations during training. Experiments demonstrate that ARS consistently improves detection and achieves substantial gains over strong baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on10
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMsMiao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li et al.ICLR 2024 · 867 citations
- Uncertainty Estimation in Autoregressive Structured PredictionAndrey Malinin, Mark J. F. GalesICLR 2021 · 439 citations
- INSIDE: LLMs' Internal States Retain the Power of Hallucination DetectionChao Chen, Kai Liu, Ze Chen, Yi Gu et al.ICLR 2024 · 281 citations
Related papers
- Joint Evaluation of Answer and Reasoning Consistency for Hallucination Detection in Large Reasoning ModelsChangyue Wang, Weihang Su, Qingyao Ai, Yiqun LiuAAAI 2026 · 13 citations
- RFS-Guard: Detecting Reasoning Hallucinations via Cross-Phase Routing Focus in Large Reasoning ModelsZihang Liu, Zhouhua Fang, Hui Liu, Zhiwei Liu et al.ACL 2026
- SHARP: Steering Hallucination in LVLMs via Representation EngineeringJunfei Wu, Yue Ding, Guofan Liu, Tianze Xia et al.EMNLP 2025
- Mind the Gap: Catching Hallucinations via Evidence Drop on the Reasoning ManifoldQunJie Chen, Yufei Chen, Xiaodong Yue, Linye LiICML 2026
- Answering the Wrong Question: Reasoning Trace Inversion for Abstention in LLMsAbinitha Gourabathina, Inkit Padhi, Manish Nagireddy, Subhajit Chaudhury et al.ACL 2026
