Won't Get Fooled Again: Answering Questions with False Premises
Shengding Hu, Yifan Luo, Huadong Wang, Xingyi Cheng, Zhiyuan Liu, Maosong Sun
Abstract
Pre-trained language models (PLMs) have shown unprecedented potential in various fields, especially as the backbones for question-answering (QA) systems. However, they tend to be easily deceived by tricky questions such as “How many eyes does the sun have?”. Such frailties of PLMs often allude to the lack of knowledge within them. In this paper, we find that the PLMs already possess the knowledge required to rebut such questions, and the key is how to activate the knowledge. To systematize this observation, we investigate the PLMs’ responses to one kind of tricky questions, i.e., the false premises questions (FPQs). We annotate a FalseQA dataset containing 2365 human-written FPQs, with the corresponding explanations for the false premises and the revised true premise questions. Using FalseQA, we discover that PLMs are capable of discriminating FPQs by fine-tuning on moderate numbers (e.g., 256) of examples. PLMs also generate reasonable explanations for the false premise, which serve as rebuttals. Further replaying a few general questions during training allows PLMs to excel on FPQs and general questions simultaneously. Our work suggests that once the rebuttal ability is stimulated, knowledge inside the PLMs can be effectively utilized to handle FPQs, which incentivizes the research on PLM-based QA systems. The FalseQA dataset and code are available at https://github.com/thunlp/FalseQA .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers19
- ULTRAFEEDBACK: Boosting Language Models with Scaled AI FeedbackGanqu Cui, Lifan Yuan, Ning Ding, Guanming Yao et al.ICML 2024 · 286 citations
- Preference Leakage: A Contamination Problem in LLM-as-a-judgeDawei Li, Renliang Sun, Yue Huang, Ming Zhong et al.ICLR 2026 · 150 citations
- InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward ModelingYuchun Miao, Sen Zhang, Liang Ding, Rong Bao et al.NeurIPS 2024 · 108 citations
- Don't Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM CollaborationShangbin Feng, Weijia Shi, Yike Wang, Wenxuan Ding et al.ACL 2024 · 30 citations
- SaySelf: Teaching LLMs to Express Confidence with Self-Reflective RationalesTianyang Xu, Shujin Wu, Shizhe Diao, Xiaoze Liu et al.EMNLP 2024 · 10 citations
Builds on7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat et al.ICML 2020 · 2,937 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- CREPE: Open-Domain Question Answering with False PresuppositionsXinyan Yu, Sewon Min, Luke Zettlemoyer, Hannaneh HajishirziACL 2023 · 13 citations
Related papers
- Whispers that Shake Foundations: Analyzing and Mitigating False Premise Hallucinations in Large Language ModelsHongbang Yuan, Pengfei Cao, Zhuoran Jin, Yubo Chen et al.EMNLP 2024 · 5 citations
- Towards Understanding Factual Knowledge of Large Language ModelsXuming Hu, Junzhe Chen, Xiaochuan Li, Yufei Guo et al.ICLR 2024 · 21 citations
- Pre-training Language Models with Deterministic Factual KnowledgeShaobo Li, Xiaoguang Li, Lifeng Shang, Chengjie Sun et al.EMNLP 2022 · 13 citations
- Identifying and Answering Questions with False Assumptions: An Interpretable ApproachZijie Wang, Eduardo BlancoEMNLP 2025
- A Glitch in the Matrix? Locating and Detecting Language Model Grounding with FakepediaGiovanni Monea, Maxime Peyrard, Martin Josifoski, Vishrav Chaudhary et al.ACL 2024 · 2 citations
