Reasoning Models Hallucinate More: Factuality-Aware Reinforcement Learning for Large Reasoning Models
Junyi Li, Hwee Tou Ng
Abstract
Large language models (LLMs) have significantly advanced in reasoning tasks through reinforcement learning (RL) optimization, achieving impressive capabilities across various challenging benchmarks. However, our empirical analysis reveals a critical drawback: reasoning-oriented RL fine-tuning significantly increases the prevalence of hallucinations. We theoretically analyze the RL training dynamics, identifying high-variance gradient, entropy-induced randomness, and susceptibility to spurious local optima as key factors leading to hallucinations. To address this drawback, we propose Factuality-aware Step-wise Policy Optimization (FSPO), an innovative RL fine-tuning algorithm incorporating explicit factuality verification at each reasoning step. FSPO leverages automated verification against given evidence to dynamically adjust token-level advantage values, incentivizing factual correctness throughout the reasoning process. Experiments across mathematical reasoning and hallucination benchmarks using Qwen2.5 and Llama models demonstrate that FSPO effectively reduces hallucinations while enhancing reasoning accuracy, substantially improving both reliability and performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1695ac70-2d47-428c-9a51-489dc8211b3fCited by top-tier papers6
- Thinking in Uncertainty: Mitigating Hallucinations in MLRMs with Latent Entropy-Aware DecodingZhongxing Xu, Zhonghua Wang, Zhe Qian, Dachuan Shi et al.CVPR 2026 · 16 citations
- KnowRL: Exploring Knowledgeable Reinforcement Learning for FactualityBaochang Ren, Shuofei Qiao, Ningyu Zhang, Da Zheng et al.ACL 2026 · 12 citations
- Towards a Mechanistic Understanding of Large Reasoning Models: A Survey of Training, Inference, and FailuresYi Hu, Jiaqi Gu, Ruxin Wang, Zijun Yao et al.ACL 2026 · 5 citations
- Towards Pareto-Optimal Tool-Integrated Agents with Pareto Ranking Policy OptimizationJunyi Li, Xiaowei Qian, Yingyi Zhang, Wenlin Zhang et al.ICML 2026
- MARCH: Multi-Agent Reinforced Check for HallucinationZhuo Li, Yupeng Zhang, Pengyu Cheng, Jiajun Song et al.ACL 2026
Builds on8
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu et al.ACL 2023 · 540 citations
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language ModelsPotsawee Manakul, Adian Liusie, Mark J. F. GalesEMNLP 2023 · 331 citations
- HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language ModelsJunyi Li, Xiaoxue Cheng, Xin Zhao, Jian-Yun Nie et al.EMNLP 2023 · 224 citations
Related papers
- FLAME : Factuality-Aware Alignment for Large Language ModelsSheng-Chieh Lin, Luyu Gao, Barlas Oguz, Wenhan Xiong et al.NeurIPS 2024 · 63 citations
- FAPO: Flawed-Aware Policy Optimization for Efficient and Reliable ReasoningYuyang Ding, Chi Zhang, Juntao Li, Haibin Lin et al.ICLR 2026 · 7 citations
- Learning to Reason for Hallucination Span DetectionHsuan Su, Ting-Yao Hu, Hema Swetha Koppula, Kundan Krishna et al.ICLR 2026 · 8 citations
- GPO: Learning from Critical Steps to Improve LLM ReasoningJiahao Yu, Zelei Cheng, Xian Wu, Xinyu XingNeurIPS 2025 · 10 citations
- The Reasoning Trap: How Enhancing LLM Reasoning Amplifies Tool HallucinationChenlong Yin, Zeyang Sha, Shiwen Cui, Changhua Meng et al.ACL 2026 · 7 citations
