BELLE: A Bi-Level Multi-Agent Reasoning Framework for Multi-Hop Question Answering
Taolin Zhang, Dongyang Li, Qizhou Chen, Chengyu Wang, Xiaofeng He
摘要
Multi-hop question answering (QA) involves finding multiple relevant passages and performing step-by-step reasoning to answer complex questions. Previous works on multi-hop QA employ specific methods from different modeling perspectives based on large language models (LLMs), regardless of question types. In this paper, we first conduct an in-depth analysis of public multi-hop QA benchmarks, categorizing questions into four types and evaluating five types of cutting-edge methods: Chainof-Thought (CoT), Single-step, Iterative-step, Sub-step, and Adaptive-step. We find that different types of multi-hop questions exhibit varying degrees of sensitivity to different types of methods. Thus, we propose a Bi-levEL muLti-agEnt reasoning (BELLE) framework to address multi-hop QA by specifically focusing on the correspondence between question types and methods, with each type of method regarded as an "operator" by prompting LLMs differently. The first level of BELLE includes multiple agents that debate to formulate an executable plan of combined "operators" to address the multi-hop QA task comprehensively. During the debate, in addition to the basic roles of affirmative debater, negative debater, and judge, at the second level, we further leverage fast and slow debaters to monitor whether changes in viewpoints are reasonable. Extensive experiments demonstrate that BELLE significantly outperforms strong baselines in various datasets. Additionally, the model consumption of BELLE is higher cost-effectiveness than that of single models in more complex multihop QA scenarios.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- AgentOCR: Reimagining Agent History via Optical Self-CompressionLang Feng, Fuchao Yang, Feng Chen, Xin Cheng 等ACL 2026 · 被引用 17 次
- Mnemosyne: Accelerating Multi-Hop Question Answering via Cache Hit Order FittingHaizhou Du, Jiujiu Li, Dongyang Li, Luobin Huang 等AAAI 2026
- THOR: A Theta-Gamma Hierarchical Oscillatory Reasoning Framework for Multi-hop QAZiyang Ling, Ronald X. Xu, Mingzhai SunACL 2026
- MAB-DQA: Addressing Query Aspect Importance in Document Question Answering with Multi-Armed BanditsYixin Xiang, Yunshan Ma, Xiaoyu Du, Yibing Chen 等ACL 2026
- Disco-RAG: Discourse-Aware Retrieval-Augmented GenerationDongqi Liu, Hang Ding, Qiming Feng, Xurong Xie 等ACL 2026
它引用的顶会 Paper27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model SocietyGuohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin 等NeurIPS 2023 · 被引用 1,975 次
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai 等ICML 2022 · 被引用 1,629 次
相关 Paper
- Tree-of-Reasoning Question Decomposition for Complex Question Answering with Large Language ModelsKun Zhang, Jiali Zeng, Fandong Meng, Yuanzhuo Wang 等AAAI 2024 · 被引用 14 次
- Answering Questions by Meta-Reasoning over Multiple Chains of ThoughtOri Yoran, Tomer Wolfson, Ben Bogin, Uri Katz 等EMNLP 2023 · 被引用 30 次
- ReAgent: Reversible Multi-Agent Reasoning for Knowledge-Enhanced Multi-Hop QAXinjie Zhao, Fan Gao, Xingyu Song, Yingjian Chen 等EMNLP 2025 · 被引用 1 次
- MAD-Logic: Multi-Agent Debate Enhances Symbolic Translation and ReasoningHaocheng Yang, Fengxiang Cheng, Tianjun Yao, Mengyue Yang 等ICLR 2026
- BeamAggR: Beam Aggregation Reasoning over Multi-source Knowledge for Multi-hop Question AnsweringZheng Chu, Jingchang Chen, Qianglong Chen, Haotian Wang 等ACL 2024 · 被引用 8 次
