Post-Abstention: Towards Reliably Re-Attempting the Abstained Instances in QA
Neeraj Varshney, Chitta Baral
摘要
Despite remarkable progress made in natural language processing, even the state-of-the-art models often make incorrect predictions. Such predictions hamper the reliability of systems and limit their widespread adoption in real-world applications. ‘Selective prediction’ partly addresses the above concern by enabling models to abstain from answering when their predictions are likely to be incorrect. While selective prediction is advantageous, it leaves us with a pertinent question ‘what to do after abstention’. To this end, we present an explorative study on ‘Post-Abstention’, a task that allows re-attempting the abstained instances with the aim of increasing coverage of the system without significantly sacrificing its accuracy. We first provide mathematical formulation of this task and then explore several methods to solve it. Comprehensive experiments on 11 QA datasets show that these methods lead to considerable risk improvements –performance metric of the Post-Abstention task– both in the in-domain and the out-of-domain settings. We also conduct a thorough analysis of these results which further leads to several interesting findings. Finally, we believe that our work will encourage and facilitate further research in this important area of addressing the reliability of NLP systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Mitigating Temporal Misalignment by Discarding Outdated FactsMichael J. Q. Zhang, Eunsol ChoiEMNLP 2023 · 被引用 6 次
- Teaching LLMs to Abstain across Languages via Multilingual FeedbackShangbin Feng, Weijia Shi, Yike Wang, Wenxuan Ding 等EMNLP 2024 · 被引用 4 次
- CertainlyUncertain: A Benchmark and Metric for Multimodal Epistemic and Aleatoric AwarenessKhyathi Raghavi Chandu, Linjie Li, Anas Awadalla, Ximing Lu 等ICLR 2025
它引用的顶会 Paper8
- Generating Clarifying Questions for Information RetrievalHamed Zamani, Susan T. Dumais, Nick Craswell, Paul N. Bennett 等WWW 2020 · 被引用 238 次
- Contrastive Test-Time AdaptationDian Chen, Dequan Wang, Trevor Darrell, Sayna EbrahimiCVPR 2022 · 被引用 219 次
- Building and Evaluating Open-Domain Dialogue Corpora with Clarifying QuestionsMohammad Aliannejadi, Julia Kiseleva, Aleksandr Chuklin, Jeff Dalton 等EMNLP 2021 · 被引用 61 次
- ILDAE: Instance-Level Difficulty Analysis of Evaluation DataNeeraj Varshney, Swaroop Mishra, Chitta BaralACL 2022 · 被引用 21 次
- Dataset Cartography: Mapping and Diagnosing Datasets with Training DynamicsSwabha Swayamdipta, Roy Schwartz, Nicholas Lourie, Yizhong Wang 等EMNLP 2020 · 被引用 12 次
相关 Paper
- Improving Selective Visual Question Answering by Learning from Your PeersCorentin Dancette, Spencer Whitehead, Rishabh Maheshwary, Ramakrishna Vedantam 等CVPR 2023
- The Art of Abstention: Selective Prediction and Error Regularization for Natural Language ProcessingJi Xin, Raphael Tang, Yaoliang Yu, Jimmy LinACL 2021
- SAFER: Risk-Constrained Sample-then-Filter in Large Language ModelsQingni Wang, Yue Fan, Xin WangICLR 2026 · 被引用 8 次
- Selective Question Answering under Domain ShiftAmita Kamath, Robin Jia, Percy LiangACL 2020 · 被引用 121 次
- Rejectors in the Wild: Deployment Barriers for LLM RejectorsErik Schönwälder, Claudio Hartmann, Wolfgang LehnerKDD 2026
