I Could've Asked That: Reformulating Unanswerable Questions
Wenting Zhao, Ge Gao, Claire Cardie, Alexander M. Rush
Abstract
When seeking information from unfamiliar documents, users frequently pose questions that cannot be answered by the documents. While existing large language models (LLMs) identify these unanswerable questions, they do not assist users in reformulating their questions, thereby reducing their overall utility. We curate COULDASK, an evaluation benchmark composed of existing and new datasets for document-grounded question answering, specifically designed to study reformulating unanswerable questions. We evaluate stateof-the-art open-source and proprietary LLMs on COULDASK. The results demonstrate the limited capabilities of these models in reformulating questions. Specifically, GPT-4 and Llama2-7B successfully reformulate questions only 26% and 12% of the time, respectively. Error analysis shows that 62% of the unsuccessful reformulations stem from the models merely rephrasing the questions or even generating identical questions. We publicly release the benchmark 1 and the code to reproduce the experiments 2 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 26270a8a-b32d-4aa0-b1c6-c75c33ebe038Builds on7
- Large language models are few-shot clinical information extractorsMonica Agrawal, Stefan Hegselmann, Hunter Lang, Yoon Kim et al.EMNLP 2022 · 285 citations
- AmbigQA: Answering Ambiguous Open-domain QuestionsSewon Min, Julian Michael, Hannaneh Hajishirzi, Luke ZettlemoyerEMNLP 2020 · 162 citations
- ClarifyDelphi: Reinforced Clarification Questions with Defeasibility Rewards for Social and Moral SituationsValentina Pyatkin, Jena D. Hwang, Vivek Srikumar, Ximing Lu et al.ACL 2023 · 12 citations
- Won't Get Fooled Again: Answering Questions with False PremisesShengding Hu, Yifan Luo, Huadong Wang, Xingyi Cheng et al.ACL 2023 · 5 citations
- Continually Improving Extractive QA via Human FeedbackGe Gao, Hung-Ting Chen, Yoav Artzi, Eunsol ChoiEMNLP 2023 · 5 citations
Related papers
- CounselBench: A Large-Scale Expert Evaluation and Adversarial Benchmarking of Large Language Models in Mental Health Question AnsweringYahan Li, Jifan Yao, John Bosco S. Bunyi, Adam C. Frank et al.ICLR 2026 · 24 citations
- NL2SQLBench: A Modular Benchmarking Framework for LLM-Enabled NL2SQL SolutionsShizheng Hou, Wenqi Pei, Nuo Chen, Quang-Trung Ta et al.VLDB 2026 · 1 citation
- RefineBench: Evaluating Refinement Capability of Language Models via ChecklistsYoung-Jun Lee, Seungone Kim, Byung-Kwan Lee, Minkyeong Moon et al.ICLR 2026 · 13 citations
- Large Language Models Struggle with Unreasonability in Math ProblemsJingyuan Ma, Damai Dai, Zihang Yuan, Rui Li et al.AAAI 2026 · 10 citations
- SQUAB: Evaluating LLM robustness to Ambiguous and Unanswerable Questions in Semantic ParsingSimone Papicchio, Luca Cagliero, Paolo PapottiEMNLP 2025
