On the Challenges of Evaluating Compositional Explanations in Multi-Hop Inference: Relevance, Completeness, and Expert Ratings
Peter Jansen, Kelly J. Smith, Dan Moreno, Huitzilin Ortiz
Abstract
Building compositional explanations requires models to combine two or more facts that, together, describe why the answer to a question is correct. Typically, these "multi-hop" explanations are evaluated relative to one (or a small number of) gold explanations. In this work, we show these evaluations substantially underestimate model performance, both in terms of the relevance of included facts, as well as the completeness of model-generated explanations, because models regularly discover and produce valid explanations that are different than gold explanations. To address this, we construct a large corpus of 126k domain-expert (science teacher) relevance ratings that augment a corpus of explanations to standardized science exam questions, discovering 80k additional relevant facts not rated as gold. We build three strong models based on different methodologies (generation, ranking, and schemas), and empirically show that while expert-augmented ratings provide better estimates of explanation quality, both original (gold) and expertaugmented automatic evaluations still substantially underestimate performance by up to 36% when compared with full manual expert judgements, with different models being disproportionately affected. This poses a significant methodological challenge to accurately evaluating explanations produced by compositional reasoning models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext de9b56cd-2cba-4b62-9ca3-b11c66e8621bCited by top-tier papers2
- ScienceWorld: Is your Agent Smarter than a 5th Grader?Ruoyao Wang, Peter A. Jansen, Marc-Alexandre Côté, Prithviraj AmmanabroluEMNLP 2022 · 1 citation
- From Models to Microtheories: Distilling a Model's Topical Knowledge for Grounded Question-AnsweringNathaniel Weir, Bhavana Dalvi Mishra, Orion Weller, Oyvind Tafjord et al.ICLR 2025
Builds on3
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
- QASC: A Dataset for Question Answering via Sentence CompositionTushar Khot, Peter Clark, Michal Guerquin, Peter Jansen et al.AAAI 2020 · 387 citations
- Learning to Explain: Datasets and Models for Identifying Valid Reasoning Chains in Multihop Question-AnsweringHarsh Jhamtani, Peter ClarkEMNLP 2020 · 2 citations
Related papers
- Hybrid Autoregressive Inference for Scalable Multi-Hop Explanation RegenerationMarco Valentino, Mokanarangan Thayaparan, Deborah Ferreira, André FreitasAAAI 2022 · 26 citations
- Answering Questions by Meta-Reasoning over Multiple Chains of ThoughtOri Yoran, Tomer Wolfson, Ben Bogin, Uri Katz et al.EMNLP 2023 · 30 citations
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu et al.NeurIPS 2022 · 2,727 citations
- QUASER: Question Answering with Scalable Extractive RationalizationAsish Ghoshal, Srinivasan Iyer, Bhargavi Paranjape, Kushal Lakhotia et al.SIGIR 2022 · 2 citations
- F1 is Not Enough! Models and Evaluation Towards User-Centered Explainable Question AnsweringHendrik Schuff, Heike Adel, Ngoc Thang VuEMNLP 2020
