F1 is Not Enough! Models and Evaluation Towards User-Centered Explainable Question Answering
Hendrik Schuff, Heike Adel, Ngoc Thang Vu
摘要
Explainable question answering systems predict an answer together with an explanation showing why the answer has been selected. The goal is to enable users to assess the correctness of the system and understand its reasoning process. However, we show that current models and evaluation settings have shortcomings regarding the coupling of answer and explanation which might cause serious issues in user experience. As a remedy, we propose a hierarchical model and a new regularization term to strengthen the answer-explanation coupling as well as two evaluation scores to quantify the coupling. We conduct experiments on the HOTPOTQA benchmark data set and perform a user study. The user study shows that our models increase the ability of the users to judge the correctness of the system and that scores like F 1 are not enough to estimate the usefulness of a model in a practical setting with human users. Our scores are better aligned with user experience, making them promising candidates for model selection.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Measuring Association Between Labels and Free-Text RationalesSarah Wiegreffe, Ana Marasovic, Noah A. SmithEMNLP 2021 · 被引用 12 次
- Reasoning over Hierarchical Question Decomposition Tree for Explainable Question AnsweringJiajie Zhang, Shulin Cao, Tingjian Zhang, Xin Lv 等ACL 2023 · 被引用 4 次
- Grow-and-Clip: Informative-yet-Concise Evidence Distillation for Answer ExplanationYuyan Chen, Yanghua Xiao, Bang LiuICDE 2022 · 被引用 1 次
- Teaching Vision-Language Models to Ask: Resolving Ambiguity in Visual QuestionsPu Jian, Donglei Yu, Wen Yang, Shuo Ren 等ACL 2025
- What's in Your Head? Emergent Behaviour in Multi-Task Transformer ModelsMor Geva, Uri Katz, Aviv Ben-Arie, Jonathan BerantEMNLP 2021
它引用的顶会 Paper6
- Learning to Retrieve Reasoning Paths over Wikipedia Graph for Question AnsweringAkari Asai, Kazuma Hashimoto, Hannaneh Hajishirzi, Richard Socher 等ICLR 2020 · 被引用 322 次
- Hierarchical Graph Network for Multi-hop Question AnsweringYuwei Fang, Siqi Sun, Zhe Gan, Rohit Pillai 等EMNLP 2020 · 被引用 157 次
- Select, Answer and Explain: Interpretable Multi-Hop Reading Comprehension over Multiple DocumentsMing Tu, Kevin Huang, Guangtao Wang, Jing Huang 等AAAI 2020 · 被引用 155 次
- Transformer-XH: Multi-Evidence Reasoning with eXtra Hop AttentionChen Zhao, Chenyan Xiong, Corby Rosset, Xia Song 等ICLR 2020 · 被引用 120 次
- Differentiable Reasoning over a Virtual Knowledge BaseBhuwan Dhingra, Manzil Zaheer, Vidhisha Balachandran, Graham Neubig 等ICLR 2020 · 被引用 91 次
相关 Paper
- Robustifying Multi-hop QA through Pseudo-Evidentiality TrainingKyungjae Lee, Seung-won Hwang, Sang-eun Han, Dohyeon LeeACL 2021
- DocVXQA: Context-Aware Visual Explanations for Document Question AnsweringMohamed Ali Souibgui, Changkyu Choi, Andrey Barsky, Kangsoo Jung 等ICML 2025
- Learning to Explain: Datasets and Models for Identifying Valid Reasoning Chains in Multihop Question-AnsweringHarsh Jhamtani, Peter ClarkEMNLP 2020 · 被引用 2 次
- QUASER: Question Answering with Scalable Extractive RationalizationAsish Ghoshal, Srinivasan Iyer, Bhargavi Paranjape, Kushal Lakhotia 等SIGIR 2022 · 被引用 2 次
- On the Challenges of Evaluating Compositional Explanations in Multi-Hop Inference: Relevance, Completeness, and Expert RatingsPeter Jansen, Kelly J. Smith, Dan Moreno, Huitzilin OrtizEMNLP 2021
