BELLE: A Bi-Level Multi-Agent Reasoning Framework for Multi-Hop Question Answering
Taolin Zhang, Dongyang Li, Qizhou Chen, Chengyu Wang, Xiaofeng He
Abstract
Multi-hop question answering (QA) involves finding multiple relevant passages and performing step-by-step reasoning to answer complex questions. Previous works on multi-hop QA employ specific methods from different modeling perspectives based on large language models (LLMs), regardless of question types. In this paper, we first conduct an in-depth analysis of public multi-hop QA benchmarks, categorizing questions into four types and evaluating five types of cutting-edge methods: Chainof-Thought (CoT), Single-step, Iterative-step, Sub-step, and Adaptive-step. We find that different types of multi-hop questions exhibit varying degrees of sensitivity to different types of methods. Thus, we propose a Bi-levEL muLti-agEnt reasoning (BELLE) framework to address multi-hop QA by specifically focusing on the correspondence between question types and methods, with each type of method regarded as an "operator" by prompting LLMs differently. The first level of BELLE includes multiple agents that debate to formulate an executable plan of combined "operators" to address the multi-hop QA task comprehensively. During the debate, in addition to the basic roles of affirmative debater, negative debater, and judge, at the second level, we further leverage fast and slow debaters to monitor whether changes in viewpoints are reasonable. Extensive experiments demonstrate that BELLE significantly outperforms strong baselines in various datasets. Additionally, the model consumption of BELLE is higher cost-effectiveness than that of single models in more complex multihop QA scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d64d407c-96f5-45df-a5fa-a53e756b5f53Cited by top-tier papers5
- AgentOCR: Reimagining Agent History via Optical Self-CompressionLang Feng, Fuchao Yang, Feng Chen, Xin Cheng et al.ACL 2026 · 17 citations
- Mnemosyne: Accelerating Multi-Hop Question Answering via Cache Hit Order FittingHaizhou Du, Jiujiu Li, Dongyang Li, Luobin Huang et al.AAAI 2026
- THOR: A Theta-Gamma Hierarchical Oscillatory Reasoning Framework for Multi-hop QAZiyang Ling, Ronald X. Xu, Mingzhai SunACL 2026
- MAB-DQA: Addressing Query Aspect Importance in Document Question Answering with Multi-Armed BanditsYixin Xiang, Yunshan Ma, Xiaoyu Du, Yibing Chen et al.ACL 2026
- Disco-RAG: Discourse-Aware Retrieval-Augmented GenerationDongqi Liu, Hang Ding, Qiming Feng, Xurong Xie et al.ACL 2026
Builds on27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model SocietyGuohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin et al.NeurIPS 2023 · 1,975 citations
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai et al.ICML 2022 · 1,629 citations
Related papers
- Tree-of-Reasoning Question Decomposition for Complex Question Answering with Large Language ModelsKun Zhang, Jiali Zeng, Fandong Meng, Yuanzhuo Wang et al.AAAI 2024 · 14 citations
- Answering Questions by Meta-Reasoning over Multiple Chains of ThoughtOri Yoran, Tomer Wolfson, Ben Bogin, Uri Katz et al.EMNLP 2023 · 30 citations
- ReAgent: Reversible Multi-Agent Reasoning for Knowledge-Enhanced Multi-Hop QAXinjie Zhao, Fan Gao, Xingyu Song, Yingjian Chen et al.EMNLP 2025 · 1 citation
- MAD-Logic: Multi-Agent Debate Enhances Symbolic Translation and ReasoningHaocheng Yang, Fengxiang Cheng, Tianjun Yao, Mengyue Yang et al.ICLR 2026
- BeamAggR: Beam Aggregation Reasoning over Multi-source Knowledge for Multi-hop Question AnsweringZheng Chu, Jingchang Chen, Qianglong Chen, Haotian Wang et al.ACL 2024 · 8 citations
