Learn to Refuse: Making Large Language Models More Controllable and Reliable through Knowledge Scope Limitation and Refusal Mechanism
Lang Cao
Abstract
Large language models (LLMs) have demonstrated impressive language understanding and generation capabilities, enabling them to answer a wide range of questions across various domains. However, these models are not flawless and often produce responses that contain errors or misinformation. These inaccuracies, commonly referred to as hallucinations, render LLMs unreliable and even unusable in many scenarios. In this paper, our focus is on mitigating the issue of hallucination in LLMs, particularly in the context of question-answering. Instead of attempting to answer all questions, we explore a refusal mechanism that instructs LLMs to refuse to answer challenging questions in order to avoid errors. We then propose a simple yet effective solution called Learn to Refuse (L2R), which incorporates the refusal mechanism to enable LLMs to recognize and refuse to answer questions that they find difficult to address. To achieve this, we utilize a structured knowledge base to represent all the LLM’s understanding of the world, enabling it to provide traceable gold knowledge. This knowledge base is separate from the LLM and initially empty. It can be filled with validated knowledge and progressively expanded. When an LLM encounters questions outside its domain, the system recognizes its knowledge scope and determines whether it can answer the question independently. Additionally, we introduce a method for automatically and efficiently expanding the knowledge base of LLMs. Through qualitative and quantitative analysis, we demonstrate that our approach enhances the controllability and reliability of LLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6591d168-cf0b-4c9e-8fdf-fcca7dd6ea02Cited by top-tier papers7
- Perception of Knowledge Boundary for Large Language Models through Semi-open-ended Question AnsweringZhihua Wen, Zhiliang Tian, Zexin Jian, Zhen Huang et al.NeurIPS 2024 · 29 citations
- TableMaster: A Recipe to Advance Table Understanding with Language ModelsLang Cao, Hanbing LiuICLR 2026 · 20 citations
- DUALBREACH: Efficient Dual-Jailbreaking via Target-Driven Initialization and Multi-Target OptimizationXinzhe Huang, Kedong Xiu, Tianhang Zheng, Churui Zeng et al.NDSS 2026 · 14 citations
- Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual TasksWenbo Pan, Jie Xu, Qiguang Chen, Junhao Dong et al.ICLR 2026 · 8 citations
- MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language ModelsShrey Pandit, Jiawei Xu, Junyuan Hong, Zhangyang Wang et al.EMNLP 2025 · 8 citations
Builds on10
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister et al.NeurIPS 2023 · 1,549 citations
Related papers
- Do Retrieval Augmented Language Models Know When They Don't Know?Youchao Zhou, Heyan Huang, Yicheng Liu, Rui Dai et al.AAAI 2026
- Detecting and Reducing the Factual Hallucinations of Large Language Models with Metamorphic TestingWeibin Wu, Yuhang Cao, Ning Yi, Rongyi Ou et al.FSE 2025 · 5 citations
- Alleviating Hallucinations from Knowledge Misalignment in Large Language Models via Selective Abstention LearningLei Huang, Xiaocheng Feng, Weitao Ma, Yuchun Fan et al.ACL 2025
- The Dawn After the Dark: An Empirical Study on Factuality Hallucination in Large Language ModelsJunyi Li, Jie Chen, Ruiyang Ren, Xiaoxue Cheng et al.ACL 2024 · 49 citations
- Attributive Reasoning for Hallucination Diagnosis of Large Language ModelsYuyan Chen, Zehao Li, Shuangjie You, Zhengyu Chen et al.AAAI 2025 · 33 citations
