FLAME : Factuality-Aware Alignment for Large Language Models
Sheng-Chieh Lin, Luyu Gao, Barlas Oguz, Wenhan Xiong, Jimmy Lin, Scott Yih, Xilun Chen
摘要
Alignment is a standard procedure to fine-tune pre-trained large language models (LLMs) to follow natural language instructions and serve as helpful AI assistants. We have observed, however, that the conventional alignment process fails to enhance the factual accuracy of LLMs, and often leads to the generation of more false facts (i.e. hallucination). In this paper, we study how to make the LLM alignment process more factual, by first identifying factors that lead to hallucination in both alignment steps: supervised fine-tuning (SFT) and reinforcement learning (RL). In particular, we find that training the LLM on new knowledge or unfamiliar texts can encourage hallucination. This makes SFT less factual as it trains on human labeled data that may be novel to the LLM. Furthermore, reward functions used in standard RL can also encourage hallucination, because it guides the LLM to provide more helpful responses on a diverse set of instructions, often preferring longer and more detailed responses. Based on these observations, we propose factuality-aware alignment, comprised of factuality-aware SFT and factuality-aware RL through direct preference optimization. Experiments show that our proposed factuality-aware alignment guides LLMs to output more factual responses while maintaining instruction-following capability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper27
- Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?Zorik Gekhman, Gal Yona, Roee Aharoni, Matan Eyal 等EMNLP 2024 · 被引用 53 次
- Parameters vs. Context: Fine-Grained Control of Knowledge Reliance in Language ModelsBaolong Bi, Shenghua Liu, Yiwei Wang, Yilong Xu 等ICLR 2026 · 被引用 47 次
- Learning to Reason for FactualityXilun Chen, Ilia Kulikov, Vincent-Pierre Berges, Barlas Oğuz 等ICML 2026 · 被引用 22 次
- LoGU: Long-form Generation with Uncertainty ExpressionsRuihan Yang, Caiqi Zhang, Zhisong Zhang, Xinting Huang 等ACL 2025 · 被引用 21 次
- KnowRL: Exploring Knowledgeable Reinforcement Learning for FactualityBaochang Ren, Shuofei Qiao, Ningyu Zhang, Da Zheng 等ACL 2026 · 被引用 12 次
它引用的顶会 Paper26
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
相关 Paper
- Self-Alignment for Factuality: Mitigating Hallucinations in LLMs via Self-EvaluationXiaoying Zhang, Baolin Peng, Ye Tian, Jingyan Zhou 等ACL 2024
- Mask-DPO: Generalizable Fine-grained Factuality Alignment of LLMsYuzhe Gu, Wenwei Zhang, Chengqi Lyu, Dahua Lin 等ICLR 2025
- Fine-Tuning Language Models for FactualityKatherine Tian, Eric Mitchell, Huaxiu Yao, Christopher D. Manning 等ICLR 2024 · 被引用 270 次
- Reasoning Models Hallucinate More: Factuality-Aware Reinforcement Learning for Large Reasoning ModelsJunyi Li, Hwee Tou NgNeurIPS 2025 · 被引用 20 次
- An Emulator for Fine-tuning Large Language Models using Small Language ModelsEric Mitchell, Rafael Rafailov, Archit Sharma, Chelsea Finn 等ICLR 2024 · 被引用 68 次
