Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?
Zorik Gekhman, Gal Yona, Roee Aharoni, Matan Eyal, Amir Feder, Roi Reichart, Jonathan Herzig
摘要
When large language models are aligned via supervised fine-tuning, they may encounter new factual information that was not acquired through pre-training. It is often conjectured that this can teach the model the behavior of hallucinating factually incorrect responses, as the model is trained to generate facts that are not grounded in its pre-existing knowledge. In this work, we study the impact of such exposure to new knowledge on the capability of the fine-tuned model to utilize its pre-existing knowledge. To this end, we design a controlled setup, focused on closedbook QA, where we vary the proportion of the fine-tuning examples that introduce new knowledge. We demonstrate that large language models struggle to acquire new factual knowledge through fine-tuning, as fine-tuning examples that introduce new knowledge are learned significantly slower than those consistent with the model's knowledge. However, we also find that as the examples with new knowledge are eventually learned, they linearly increase the model's tendency to hallucinate. Taken together, our results highlight the risk in introducing new factual knowledge through fine-tuning, and support the view that large language models mostly acquire factual knowledge through pre-training, whereas finetuning teaches them to use it more efficiently.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper37
- Alignment for HonestyYuqing Yang, Ethan Chern, Xipeng Qiu, Graham Neubig 等NeurIPS 2024 · 被引用 82 次
- Understanding Finetuning for Factual Knowledge ExtractionGaurav Rohit Ghosal, Tatsunori Hashimoto, Aditi RaghunathanICML 2024 · 被引用 36 次
- Stylus: Automatic Adapter Selection for Diffusion ModelsMichael Luo, Justin Wong, Brandon Trabucco, Yanping Huang 等NeurIPS 2024 · 被引用 27 次
- Learning to Reason for FactualityXilun Chen, Ilia Kulikov, Vincent-Pierre Berges, Barlas Oğuz 等ICML 2026 · 被引用 22 次
- RAG-Enhanced Collaborative LLM Agents for Drug DiscoveryNamkyeong Lee, Edward De Brouwer, Ehsan Hajiramezanali, Tommaso Biancalani 等AAAI 2026 · 被引用 21 次
它引用的顶会 Paper26
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
相关 Paper
- FLAME : Factuality-Aware Alignment for Large Language ModelsSheng-Chieh Lin, Luyu Gao, Barlas Oguz, Wenhan Xiong 等NeurIPS 2024 · 被引用 63 次
- Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMsOded Ovadia, Menachem Brief, Moshik Mishaeli, Oren ElishaEMNLP 2024 · 被引用 89 次
- Alleviating Hallucinations from Knowledge Misalignment in Large Language Models via Selective Abstention LearningLei Huang, Xiaocheng Feng, Weitao Ma, Yuchun Fan 等ACL 2025
- Fictitious Synthetic Data Can Improve LLM Factuality via Prerequisite LearningYujian Liu, Shiyu Chang, Tommi S. Jaakkola, Yang ZhangICLR 2025
- KnowTuning: Knowledge-aware Fine-tuning for Large Language ModelsYougang Lyu, Lingyong Yan, Shuaiqiang Wang, Haibo Shi 等EMNLP 2024 · 被引用 3 次
