Is Multi-Hop Reasoning Really Explainable? Towards Benchmarking Reasoning Interpretability
Xin Lv, Yixin Cao, Lei Hou, Juanzi Li, Zhiyuan Liu, Yichi Zhang, Zelin Dai
Abstract
Multi-hop reasoning has been widely studied in recent years to obtain more interpretable link prediction. However, we find in experiments that many paths given by these models are actually unreasonable, while little work has been done on interpretability evaluation for them. In this paper, we propose a unified framework to quantitatively evaluate the interpretability of multi-hop reasoning models so as to advance their development. In specific, we define three metrics, including path recall, local interpretability, and global interpretability for evaluation, and design an approximate strategy to calculate these metrics using the interpretability scores of rules. Furthermore, we manually annotate all possible rules and establish a Benchmark to detect the Interpretability of Multi-hop Reasoning (BIMR). In experiments, we verify the effectiveness of our benchmark. Besides, we run nine representative baselines on our benchmark, and the experimental results show that the interpretability of current multi-hop reasoning models is less satisfactory and is 51.7% lower than the upper bound given by our benchmark. Moreover, the rule-based models outperform the multi-hop reasoning models in terms of performance and interpretability, which points to a direction for future research, i.e., how to better incorporate rule information into the multihop reasoning model. Our codes and datasets can be obtained from https://github . com/THU-KEG/BIMR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- SQUIRE: A Sequence-to-sequence Framework for Multi-hop Knowledge Graph ReasoningYushi Bai, Xin Lv, Juanzi Li, Lei Hou et al.EMNLP 2022 · 19 citations
- Unveiling Internal Reasoning Modes in LLMs: A Deep Dive into Latent Reasoning vs. Factual Shortcuts with Attribute Rate RatioYiran Yang, Haifeng Sun, Jingyu Wang, Qi Qi et al.EMNLP 2025
Builds on4
- Learning Reasoning Strategies in End-to-End Differentiable ProvingPasquale Minervini, Sebastian Riedel, Pontus Stenetorp, Edward Grefenstette et al.ICML 2020 · 102 citations
- Reasoning on Knowledge Graphs with Debate DynamicsMarcel Hildebrandt, Jorge Andres Quintero Serna, Yunpu Ma, Martin Ringsquandl et al.AAAI 2020 · 59 citations
- Dynamic Anticipation and Completion for Multi-Hop Reasoning over Sparse Knowledge GraphXin Lv, Xu Han, Lei Hou, Juanzi Li et al.EMNLP 2020 · 57 citations
- Learning to Explain: Datasets and Models for Identifying Valid Reasoning Chains in Multihop Question-AnsweringHarsh Jhamtani, Peter ClarkEMNLP 2020 · 2 citations
Related papers
- TFRBench: A Reasoning Benchmark for Evaluating Forecasting SystemsMd Atik Ahamed, Mihir Parmar, Palash Goyal, Yiwen Song et al.ICML 2026
- eXpath: Explaining Knowledge Graph Link Prediction with Ontological Closed Path RulesYe Sun, Lei Shi, Yongxin TongVLDB 2025 · 3 citations
- Seemingly Plausible Distractors in Multi-Hop Reasoning: Are Large Language Models Attentive Readers?Neeladri Bhuiya, Viktor Schlegel, Stefan WinklerEMNLP 2024 · 2 citations
- ERASER: A Benchmark to Evaluate Rationalized NLP ModelsJay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric P. Lehman et al.ACL 2020 · 36 citations
- Reasoning Runtime Behavior of a Program with LLM: How Far are We?Junkai Chen, Zhiyuan Pan, Xing Hu, Zhenhao Li et al.ICSE 2025 · 5 citations
