F1 is Not Enough! Models and Evaluation Towards User-Centered Explainable Question Answering
Hendrik Schuff, Heike Adel, Ngoc Thang Vu
Abstract
Explainable question answering systems predict an answer together with an explanation showing why the answer has been selected. The goal is to enable users to assess the correctness of the system and understand its reasoning process. However, we show that current models and evaluation settings have shortcomings regarding the coupling of answer and explanation which might cause serious issues in user experience. As a remedy, we propose a hierarchical model and a new regularization term to strengthen the answer-explanation coupling as well as two evaluation scores to quantify the coupling. We conduct experiments on the HOTPOTQA benchmark data set and perform a user study. The user study shows that our models increase the ability of the users to judge the correctness of the system and that scores like F 1 are not enough to estimate the usefulness of a model in a practical setting with human users. Our scores are better aligned with user experience, making them promising candidates for model selection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7d579c66-e73e-4555-9485-64a22516f6a2Cited by top-tier papers5
- Measuring Association Between Labels and Free-Text RationalesSarah Wiegreffe, Ana Marasovic, Noah A. SmithEMNLP 2021 · 12 citations
- Reasoning over Hierarchical Question Decomposition Tree for Explainable Question AnsweringJiajie Zhang, Shulin Cao, Tingjian Zhang, Xin Lv et al.ACL 2023 · 4 citations
- Grow-and-Clip: Informative-yet-Concise Evidence Distillation for Answer ExplanationYuyan Chen, Yanghua Xiao, Bang LiuICDE 2022 · 1 citation
- Teaching Vision-Language Models to Ask: Resolving Ambiguity in Visual QuestionsPu Jian, Donglei Yu, Wen Yang, Shuo Ren et al.ACL 2025
- What's in Your Head? Emergent Behaviour in Multi-Task Transformer ModelsMor Geva, Uri Katz, Aviv Ben-Arie, Jonathan BerantEMNLP 2021
Builds on6
- Learning to Retrieve Reasoning Paths over Wikipedia Graph for Question AnsweringAkari Asai, Kazuma Hashimoto, Hannaneh Hajishirzi, Richard Socher et al.ICLR 2020 · 322 citations
- Hierarchical Graph Network for Multi-hop Question AnsweringYuwei Fang, Siqi Sun, Zhe Gan, Rohit Pillai et al.EMNLP 2020 · 157 citations
- Select, Answer and Explain: Interpretable Multi-Hop Reading Comprehension over Multiple DocumentsMing Tu, Kevin Huang, Guangtao Wang, Jing Huang et al.AAAI 2020 · 155 citations
- Transformer-XH: Multi-Evidence Reasoning with eXtra Hop AttentionChen Zhao, Chenyan Xiong, Corby Rosset, Xia Song et al.ICLR 2020 · 120 citations
- Differentiable Reasoning over a Virtual Knowledge BaseBhuwan Dhingra, Manzil Zaheer, Vidhisha Balachandran, Graham Neubig et al.ICLR 2020 · 91 citations
Related papers
- Robustifying Multi-hop QA through Pseudo-Evidentiality TrainingKyungjae Lee, Seung-won Hwang, Sang-eun Han, Dohyeon LeeACL 2021
- DocVXQA: Context-Aware Visual Explanations for Document Question AnsweringMohamed Ali Souibgui, Changkyu Choi, Andrey Barsky, Kangsoo Jung et al.ICML 2025
- Learning to Explain: Datasets and Models for Identifying Valid Reasoning Chains in Multihop Question-AnsweringHarsh Jhamtani, Peter ClarkEMNLP 2020 · 2 citations
- QUASER: Question Answering with Scalable Extractive RationalizationAsish Ghoshal, Srinivasan Iyer, Bhargavi Paranjape, Kushal Lakhotia et al.SIGIR 2022 · 2 citations
- On the Challenges of Evaluating Compositional Explanations in Multi-Hop Inference: Relevance, Completeness, and Expert RatingsPeter Jansen, Kelly J. Smith, Dan Moreno, Huitzilin OrtizEMNLP 2021
