Learning to Reason for Factuality
Xilun Chen, Ilia Kulikov, Vincent-Pierre Berges, Barlas Oğuz, Rulin Shao, Gargi Ghosh, Scott Yih
摘要
Reasoning Large Language Models (R-LLMs) have significantly advanced complex reasoning tasks but often struggle with factuality, generating substantially more hallucinations than their non-reasoning counterparts on long-form factuality benchmarks. However, extending online Reinforcement Learning (RL), a key component in recent R-LLM advancements, to the long-form factuality setting poses several unique challenges due to the lack of reliable verification methods. Previous work has utilized automatic factuality evaluation frameworks such as FActScore to curate preference data in the offline RL setting, yet we find that directly leveraging such methods as the reward in online RL leads to reward hacking in multiple ways, such as producing less detailed or relevant responses. We propose a novel reward function that simultaneously considers the factual precision, response detail level, and answer relevance, and applies online RL to learn high quality factual reasoning. Evaluated on six long-form factuality benchmarks, our factual reasoning model achieves an average reduction of 23.1 percentage points in hallucination rate, a 23% increase in answer detail level, and no degradation in the overall response helpfulness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Reinforced Informativeness Optimization for Long-Form Retrieval-Augmented GenerationYuhao Wang, Ruiyang Ren, Yucheng Wang, Xin Zhao 等ACL 2026 · 被引用 4 次
- Train for Truth, Keep the Skills: Binary Retrieval-Augmented Reward Mitigates HallucinationsTong Chen, Akari Asai, Luke Zettlemoyer, Hannaneh Hajishirzi 等ICML 2026
- TruthRL: Incentivizing Truthful LLMs via Reinforcement LearningZhepei Wei, Xiao Yang, Kai Sun, Jiaqi Wang 等ICML 2026
- From Proof to Program: Characterizing Tool-Induced Reasoning Hallucinations in Large Language ModelsFarima Fatahi Bayat, Pouya Pezeshkpour, Estevam HruschkaACL 2026
它引用的顶会 Paper17
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human FeedbackYann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang 等NeurIPS 2023 · 被引用 948 次
- WildChat: 1M ChatGPT Interaction Logs in the WildWenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie 等ICLR 2024 · 被引用 504 次
相关 Paper
- KnowRL: Exploring Knowledgeable Reinforcement Learning for FactualityBaochang Ren, Shuofei Qiao, Ningyu Zhang, Da Zheng 等ACL 2026 · 被引用 12 次
- FLAME : Factuality-Aware Alignment for Large Language ModelsSheng-Chieh Lin, Luyu Gao, Barlas Oguz, Wenhan Xiong 等NeurIPS 2024 · 被引用 63 次
- Improving Factuality with Explicit Working MemoryMingda Chen, Yang Li, Karthik Padthe, Rulin Shao 等ACL 2025
- Reasoning Models Hallucinate More: Factuality-Aware Reinforcement Learning for Large Reasoning ModelsJunyi Li, Hwee Tou NgNeurIPS 2025 · 被引用 20 次
- Learning to Reason for Hallucination Span DetectionHsuan Su, Ting-Yao Hu, Hema Swetha Koppula, Kundan Krishna 等ICLR 2026 · 被引用 8 次
