Utilize the Flow Before Stepping into the Same River Twice: Certainty Represented Knowledge Flow for Refusal-Aware Instruction Tuning
Runchuan Zhu, Zhipeng Ma, Jiang Wu, Junyuan Gao, Jiaqi Wang, Dahua Lin, Conghui He
Abstract
Refusal-Aware Instruction Tuning (RAIT) enables Large Language Models (LLMs) to refuse to answer unknown questions. By modifying responses of unknown questions in the training data to refusal responses such as ''I don't know", RAIT enhances the reliability of LLMs and reduces their hallucination. Generally, RAIT modifies training samples based on the correctness of the initial LLM's response. However, this crude approach can cause LLMs to excessively refuse answering questions they could have correctly answered, the problem we call over-refusal. In this paper, we explore two primary causes of over-refusal: Static conflict occurs when similar samples within the LLM’s feature space receive differing supervision signals (original vs. modified ''I don't know"). Dynamic conflict arises as the LLM's evolving knowledge during SFT enables it to answer previously unanswerable questions, but the now-answerable training samples still retain the original ''I don't know" supervision signals from the initial LLM state, leading to inconsistencies. These conflicts cause the trained LLM to misclassify known questions as unknown, resulting in over-refusal. To address this issue, we introduce Certainty Represented Knowledge Flow for Refusal-Aware Instructions Tuning (CRaFT). CRaFT centers on two main contributions: First, we additionally incorporate response certainty to selectively filter and modify data, reducing static conflicts. Second, we implement preliminary rehearsal training to characterize changes in the LLM's knowledge state, which helps mitigate dynamic conflicts during the fine-tuning process. We conducted extensive experiments on open-ended question answering and multiple-choice question task. Experiment results show that CRaFT can improve LLM's overall performance during the RAIT process.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Knowledge Boundary of Large Language Models: A SurveyMoxin Li, Yong Zhao, Wenxuan Zhang, Shuaiyi Li et al.ACL 2025 · 33 citations
- Large Language Models Struggle with Unreasonability in Math ProblemsJingyuan Ma, Damai Dai, Zihang Yuan, Rui Li et al.AAAI 2026 · 10 citations
- Trust Within? Seek Beyond? Knowledge Boundary Aware Policy Optimization for Agentic SearchTao Feng, Xinke Jiang, Xinyan Hu, Yonggang Zhang et al.ACL 2026 · 1 citation
- Do Retrieval Augmented Language Models Know When They Don't Know?Youchao Zhou, Heyan Huang, Yicheng Liu, Rui Dai et al.AAAI 2026
Builds on11
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister et al.NeurIPS 2023 · 1,549 citations
- DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language ModelsYung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim et al.ICLR 2024 · 354 citations
Related papers
- Alleviating Hallucinations from Knowledge Misalignment in Large Language Models via Selective Abstention LearningLei Huang, Xiaocheng Feng, Weitao Ma, Yuchun Fan et al.ACL 2025
- High Accuracy, Less Talk (HALT): Reliable LLMs through Capability-Aligned FinetuningTim Franzmeyer, Archie Sravankumar, Lijuan Liu, Yuning Mao et al.ICLR 2026 · 2 citations
- Fictitious Synthetic Data Can Improve LLM Factuality via Prerequisite LearningYujian Liu, Shiyu Chang, Tommi S. Jaakkola, Yang ZhangICLR 2025
- Don't Half-listen: Capturing Key-part Information in Continual Instruction TuningYongquan He, Wenyuan Zhang, Xuancheng Huang, Peng Zhang et al.ACL 2025
- When Silence Is Golden: Can LLMs Learn to Abstain in Temporal QA and Beyond?Xinyu Zhou, Chang Jin, Carsten Eickhoff, Zhijiang Guo et al.ICLR 2026 · 6 citations
