Don't Just Say "I don't know"! Self-aligning Large Language Models for Responding to Unknown Questions with Explanations
Yang Deng, Yong Zhao, Moxin Li, See-Kiong Ng, Tat-Seng Chua
摘要
Despite the remarkable abilities of Large Language Models (LLMs) to answer questions, they often display a considerable level of overconfidence even when the question does not have a definitive answer. To avoid providing hallucinated answers to these unknown questions, existing studies typically investigate approaches to refusing to answer these questions. In this work, we propose a novel and scalable self-alignment method to utilize the LLM itself to enhance its response-ability to different types of unknown questions, being capable of not just refusing to answer but further proactively providing explanations to the unanswerability of unknown questions. Specifically, the Self-Align method first employ a two-stage classaware self-augmentation approach to generate a large amount of unknown question-response data. Then we conduct disparity-driven selfcuration to select qualified data for fine-tuning the LLM itself for aligning the responses to unknown questions as desired. Experimental results on two datasets across four types of unknown questions validate the superiority of the Self-Aligned method over existing baselines in terms of three types of task formulation. 1 * Equal contribution. 1 The data and code will be released at https://github. com/zhaoy777/KUQP-Dataset . Q: What animal can be found at the top of the men's Wimbledon trophy? Direct Answer A: The animal that can be found at the top of the men's Wimbledon trophy is a falcon. Unknown Question Detection A: The answer is unknown. A: The question is incorrect. Unknown Question Classification A: The question is incorrect because the Wimbledon men's singles trophy does not feature an animal at the top. Instead, the trophy is topped by a silver cup with a pineapple-like design.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Knowledge Boundary of Large Language Models: A SurveyMoxin Li, Yong Zhao, Wenxuan Zhang, Shuaiyi Li 等ACL 2025 · 被引用 33 次
- Mitigating Lost in Multi-turn Conversation via Curriculum RL with Verifiable Accuracy and Abstention RewardsMing Li, Pei Chen, Zhenhao Zhang, Tao Yang 等ACL 2026 · 被引用 3 次
- Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model UncertaintyJingyi Ren, Ante Wang, Yunghwei Lai, Xiaolong Wang 等ACL 2026 · 被引用 1 次
- Do Retrieval Augmented Language Models Know When They Don't Know?Youchao Zhou, Heyan Huang, Yicheng Liu, Rui Dai 等AAAI 2026
它引用的顶会 Paper17
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil 等ICLR 2024 · 被引用 1,798 次
相关 Paper
- UAlign: Leveraging Uncertainty Estimations for Factuality Alignment on Large Language ModelsBoyang Xue, Fei Mi, Qi Zhu, Hongru Wang 等ACL 2025 · 被引用 10 次
- Self-Alignment for Factuality: Mitigating Hallucinations in LLMs via Self-EvaluationXiaoying Zhang, Baolin Peng, Ye Tian, Jingyan Zhou 等ACL 2024
- Dynamic Rewarding with Prompt Optimization Enables Tuning-free Self-Alignment of Language ModelsSomanshu Singla, Zhen Wang, Tianyang Liu, Abdullah Ashfaq 等EMNLP 2024 · 被引用 1 次
- Knowledge Verification to Nip Hallucination in the BudFanqi Wan, Xinting Huang, Leyang Cui, Xiaojun Quan 等EMNLP 2024 · 被引用 8 次
- Can AI Assistants Know What They Don't Know?Qinyuan Cheng, Tianxiang Sun, Xiangyang Liu, Wenwei Zhang 等ICML 2024 · 被引用 48 次
