Is Multi-Hop Reasoning Really Explainable? Towards Benchmarking Reasoning Interpretability
Xin Lv, Yixin Cao, Lei Hou, Juanzi Li, Zhiyuan Liu, Yichi Zhang, Zelin Dai
摘要
Multi-hop reasoning has been widely studied in recent years to obtain more interpretable link prediction. However, we find in experiments that many paths given by these models are actually unreasonable, while little work has been done on interpretability evaluation for them. In this paper, we propose a unified framework to quantitatively evaluate the interpretability of multi-hop reasoning models so as to advance their development. In specific, we define three metrics, including path recall, local interpretability, and global interpretability for evaluation, and design an approximate strategy to calculate these metrics using the interpretability scores of rules. Furthermore, we manually annotate all possible rules and establish a Benchmark to detect the Interpretability of Multi-hop Reasoning (BIMR). In experiments, we verify the effectiveness of our benchmark. Besides, we run nine representative baselines on our benchmark, and the experimental results show that the interpretability of current multi-hop reasoning models is less satisfactory and is 51.7% lower than the upper bound given by our benchmark. Moreover, the rule-based models outperform the multi-hop reasoning models in terms of performance and interpretability, which points to a direction for future research, i.e., how to better incorporate rule information into the multihop reasoning model. Our codes and datasets can be obtained from https://github . com/THU-KEG/BIMR.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- SQUIRE: A Sequence-to-sequence Framework for Multi-hop Knowledge Graph ReasoningYushi Bai, Xin Lv, Juanzi Li, Lei Hou 等EMNLP 2022 · 被引用 19 次
- Unveiling Internal Reasoning Modes in LLMs: A Deep Dive into Latent Reasoning vs. Factual Shortcuts with Attribute Rate RatioYiran Yang, Haifeng Sun, Jingyu Wang, Qi Qi 等EMNLP 2025
它引用的顶会 Paper4
- Learning Reasoning Strategies in End-to-End Differentiable ProvingPasquale Minervini, Sebastian Riedel, Pontus Stenetorp, Edward Grefenstette 等ICML 2020 · 被引用 102 次
- Reasoning on Knowledge Graphs with Debate DynamicsMarcel Hildebrandt, Jorge Andres Quintero Serna, Yunpu Ma, Martin Ringsquandl 等AAAI 2020 · 被引用 59 次
- Dynamic Anticipation and Completion for Multi-Hop Reasoning over Sparse Knowledge GraphXin Lv, Xu Han, Lei Hou, Juanzi Li 等EMNLP 2020 · 被引用 57 次
- Learning to Explain: Datasets and Models for Identifying Valid Reasoning Chains in Multihop Question-AnsweringHarsh Jhamtani, Peter ClarkEMNLP 2020 · 被引用 2 次
相关 Paper
- TFRBench: A Reasoning Benchmark for Evaluating Forecasting SystemsMd Atik Ahamed, Mihir Parmar, Palash Goyal, Yiwen Song 等ICML 2026
- eXpath: Explaining Knowledge Graph Link Prediction with Ontological Closed Path RulesYe Sun, Lei Shi, Yongxin TongVLDB 2025 · 被引用 3 次
- Seemingly Plausible Distractors in Multi-Hop Reasoning: Are Large Language Models Attentive Readers?Neeladri Bhuiya, Viktor Schlegel, Stefan WinklerEMNLP 2024 · 被引用 2 次
- ERASER: A Benchmark to Evaluate Rationalized NLP ModelsJay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric P. Lehman 等ACL 2020 · 被引用 36 次
- Reasoning Runtime Behavior of a Program with LLM: How Far are We?Junkai Chen, Zhiyuan Pan, Xing Hu, Zhenhao Li 等ICSE 2025 · 被引用 5 次
