(QA)²: Question Answering with Questionable Assumptions
Najoung Kim, Phu Mon Htut, Samuel R. Bowman, Jackson Petty
Abstract
Naturally occurring information-seeking questions often contain questionable assumptions-assumptions that are false or unverifiable. Questions containing questionable assumptions are challenging because they require a distinct answer strategy that deviates from typical answers for information-seeking questions. For instance, the question When did Marie Curie discover Uranium? cannot be answered as a typical when question without addressing the false assumption Marie Curie discovered Uranium. In this work, we propose (QA) 2 (Question Answering with Questionable Assumptions), an open-domain evaluation dataset consisting of naturally occurring search engine queries that may or may not contain questionable assumptions. To be successful on (QA) 2 , systems must be able to detect questionable assumptions and also be able to produce adequate responses for both typical information-seeking questions and ones with questionable assumptions. Through human rater acceptability on end-to-end QA with (QA) 2 , we find that current models do struggle with handling questionable assumptions, leaving substantial headroom for progress. * Equal contribution, corresponding authors ∆ Work partly done at NYU before joining BU. δ Work done at NYU before joining Amazon. 1 We use the term questionable assumptions instead of presupposition failure to capture failures of both true presup-
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3a05556c-58b1-4633-b6f6-3c9a712a2e29Cited by top-tier papers12
- Measuring Chain of Thought Faithfulness by Unlearning Reasoning StepsMartin Tutek, Fateme Hashemi Chaleshtori, Ana Marasovic, Yonatan BelinkovEMNLP 2025 · 37 citations
- Towards Understanding Factual Knowledge of Large Language ModelsXuming Hu, Junzhe Chen, Xiaochuan Li, Yufei Guo et al.ICLR 2024 · 21 citations
- EvolveBench: A Comprehensive Benchmark for Assessing Temporal Awareness in LLMs on Evolving KnowledgeZhiyuan Zhu, Yusheng Liao, Zhe Chen, Yuhao Wang et al.ACL 2025 · 10 citations
- No Questions are Stupid, but some are Poorly Posed: Understanding Poorly-Posed Information-Seeking QuestionsNeha Srikanth, Rachel Rudinger, Jordan Lee Boyd-GraberACL 2025 · 6 citations
- Whispers that Shake Foundations: Analyzing and Mitigating False Premise Hallucinations in Large Language ModelsHongbang Yuan, Pengfei Cao, Zhuoran Jin, Yubo Chen et al.EMNLP 2024 · 5 citations
Builds on11
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat et al.ICML 2020 · 2,937 citations
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
- Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context LearningHaokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta et al.NeurIPS 2022 · 1,483 citations
- Decomposed Prompting: A Modular Approach for Solving Complex TasksTushar Khot, Harsh Trivedi, Matthew Finlayson, Yao Fu et al.ICLR 2023 · 94 citations
Related papers
- Which Linguist Invented the Lightbulb? Presupposition Verification for Question-AnsweringNajoung Kim, Ellie Pavlick, Burcu Karagol Ayan, Deepak RamachandranACL 2021
- CREPE: Open-Domain Question Answering with False PresuppositionsXinyan Yu, Sewon Min, Luke Zettlemoyer, Hannaneh HajishirziACL 2023 · 13 citations
- Identifying and Answering Questions with False Assumptions: An Interpretable ApproachZijie Wang, Eduardo BlancoEMNLP 2025
- Challenges in Information-Seeking QA: Unanswerable Questions and Paragraph RetrievalAkari Asai, Eunsol ChoiACL 2021
- IfQA: A Dataset for Open-domain Question Answering under Counterfactual PresuppositionsWenhao Yu, Meng Jiang, Peter Clark, Ashish SabharwalEMNLP 2023 · 6 citations
