Post-Abstention: Towards Reliably Re-Attempting the Abstained Instances in QA
Neeraj Varshney, Chitta Baral
Abstract
Despite remarkable progress made in natural language processing, even the state-of-the-art models often make incorrect predictions. Such predictions hamper the reliability of systems and limit their widespread adoption in real-world applications. ‘Selective prediction’ partly addresses the above concern by enabling models to abstain from answering when their predictions are likely to be incorrect. While selective prediction is advantageous, it leaves us with a pertinent question ‘what to do after abstention’. To this end, we present an explorative study on ‘Post-Abstention’, a task that allows re-attempting the abstained instances with the aim of increasing coverage of the system without significantly sacrificing its accuracy. We first provide mathematical formulation of this task and then explore several methods to solve it. Comprehensive experiments on 11 QA datasets show that these methods lead to considerable risk improvements –performance metric of the Post-Abstention task– both in the in-domain and the out-of-domain settings. We also conduct a thorough analysis of these results which further leads to several interesting findings. Finally, we believe that our work will encourage and facilitate further research in this important area of addressing the reliability of NLP systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2ab4afa0-2434-4609-a534-53dc2a5db004Cited by top-tier papers3
- Mitigating Temporal Misalignment by Discarding Outdated FactsMichael J. Q. Zhang, Eunsol ChoiEMNLP 2023 · 6 citations
- Teaching LLMs to Abstain across Languages via Multilingual FeedbackShangbin Feng, Weijia Shi, Yike Wang, Wenxuan Ding et al.EMNLP 2024 · 4 citations
- CertainlyUncertain: A Benchmark and Metric for Multimodal Epistemic and Aleatoric AwarenessKhyathi Raghavi Chandu, Linjie Li, Anas Awadalla, Ximing Lu et al.ICLR 2025
Builds on8
- Generating Clarifying Questions for Information RetrievalHamed Zamani, Susan T. Dumais, Nick Craswell, Paul N. Bennett et al.WWW 2020 · 238 citations
- Contrastive Test-Time AdaptationDian Chen, Dequan Wang, Trevor Darrell, Sayna EbrahimiCVPR 2022 · 219 citations
- Building and Evaluating Open-Domain Dialogue Corpora with Clarifying QuestionsMohammad Aliannejadi, Julia Kiseleva, Aleksandr Chuklin, Jeff Dalton et al.EMNLP 2021 · 61 citations
- ILDAE: Instance-Level Difficulty Analysis of Evaluation DataNeeraj Varshney, Swaroop Mishra, Chitta BaralACL 2022 · 21 citations
- Dataset Cartography: Mapping and Diagnosing Datasets with Training DynamicsSwabha Swayamdipta, Roy Schwartz, Nicholas Lourie, Yizhong Wang et al.EMNLP 2020 · 12 citations
Related papers
- Improving Selective Visual Question Answering by Learning from Your PeersCorentin Dancette, Spencer Whitehead, Rishabh Maheshwary, Ramakrishna Vedantam et al.CVPR 2023
- The Art of Abstention: Selective Prediction and Error Regularization for Natural Language ProcessingJi Xin, Raphael Tang, Yaoliang Yu, Jimmy LinACL 2021
- SAFER: Risk-Constrained Sample-then-Filter in Large Language ModelsQingni Wang, Yue Fan, Xin WangICLR 2026 · 8 citations
- Selective Question Answering under Domain ShiftAmita Kamath, Robin Jia, Percy LiangACL 2020 · 121 citations
- Rejectors in the Wild: Deployment Barriers for LLM RejectorsErik Schönwälder, Claudio Hartmann, Wolfgang LehnerKDD 2026
