Beyond the Answer: Advancing Multi-Hop QA with Fine-Grained Graph Reasoning and Evaluation
Qichuan Liu, Chentao Zhang, Chenfeng Zheng, Guosheng Hu, Xiaodong Li, Zhihong Zhang
Abstract
Recent advancements in large language models (LLMs) have significantly improved the performance of multi-hop question answering (MHQA) systems. Despite the success of MHQA systems, the evaluation of MHQA is not deeply investigated. Existing evaluations mainly focus on comparing the final answers of the reasoning method and given ground-truths. We argue that the reasoning process should also be evaluated because wrong reasoning process can also lead to the correct final answers. Motivated by this, we propose a "Planner-Executor-Reasoner" (PER) architecture, which forms the core of the Plan-anchored Data Preprocessing (PER-DP) and the Plan-guided Multi-Hop QA (PER-QA). The former provides the groundtruth of intermediate reasoning steps and final answers, and the latter offers them of a reasoning method. Moreover, we design a finegrained evaluation metric called Plan-aligned Stepwise Evaluation (PSE), which evaluates the intermediate reasoning steps from two aspects: planning and solving. Extensive experiments on ten types of questions demonstrate competitive reasoning performance, improved explainability of the MHQA system, and uncover issues such as "fortuitous reasoning continuance" and "latent reasoning suspension" in RAG-based MHQA systems. Besides, we also demonstrate the potential of our approach in data contamination scenarios. Our data and code have been released at https://github.com/GenIRAG/PER-PSE .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 517c4f03-8a37-474f-ba8c-fa70a662a4ffBuilds on12
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Least-to-Most Prompting Enables Complex Reasoning in Large Language ModelsDenny Zhou, Nathanael Schärli, Le Hou, Jason Wei et al.ICLR 2023 · 318 citations
- Precise Zero-Shot Dense Retrieval without Relevance LabelsLuyu Gao, Xueguang Ma, Jimmy Lin, Jamie CallanACL 2023 · 211 citations
- Query Rewriting in Retrieval-Augmented Large Language ModelsXinbei Ma, Yeyun Gong, Pengcheng He, Hai Zhao et al.EMNLP 2023 · 191 citations
- Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step QuestionsHarsh Trivedi, Niranjan Balasubramanian, Tushar Khot, Ashish SabharwalACL 2023 · 187 citations
Related papers
- MC-Search: Evaluating and Enhancing Multimodal Agentic Search with Structured Long Reasoning ChainsXuying Ning, Dongqi Fu, Tianxin Wei, Mengting Ai et al.ICLR 2026 · 14 citations
- STRIDE: Strategic Iterative Decision-Making for Retrieval-Augmented Multi-Hop Question AnsweringWei Chen, Lili Zhao, Zhi Zheng, Huijun Hou et al.SIGIR 2026
- CofCA: A STEP-WISE Counterfactual Multi-hop QA benchmarkJian Wu, Linyi Yang, Zhen Wang, Manabu Okumura et al.ICLR 2025
- REAP: Enhancing RAG with Recursive Evaluation and Adaptive Planning for Multi-Hop Question AnsweringYijie Zhu, Haojie Zhou, Wanting Hong, Tailin Liu et al.AAAI 2026
- DTKG: Dual-Track Knowledge Graph-Verified Reasoning Framework for Multi-Hop QAChanghao Wang, Yanfang Liu, Xinxin Fan, Ao Tian et al.ICML 2026
