Eigen-Agent: Adaptive Multi-Agent Scientific Reasoning with Monitor-Based RAG
Xiangru Tang, Wanghan Xu, Yujie Wang, Zijie Guo, Daniel Shao, Cixuan Zhang, Ziyi Wang, Lixin Zhang, Frank Wan, Zhenfei Yin, Wenlong Zhang, Lei Bai
Abstract
Large language models (LLMs) have recently shown strong progress on scientific reasoning, yet two major bottlenecks remain. First, explicit retrieval fragments reasoning, imposing a hidden "tool tax" of extra tokens and steps. Second, multiagent pipelines often dilute strong solutions by averaging across all candidates. We address these challenges with a unified framework that combines implicit retrieval and structured collaboration. At its foundation, a Monitor-based retrieval module operates at the token level, integrating external knowledge with minimal disruption to reasoning. On top of this substrate, Hierarchical Solution Refinement (HSR) iteratively designates each candidate as an anchor to be repaired by its peers, while Quality-Aware Iterative Reasoning (QAIR) adapts refinement to solution quality. On Humanitys Last Exam (HLE) Bio/Chem Gold, our framework achieves 48.3% accuracy-the highest reported to date, surpassing the strongest agent baseline by 13.4 points and leading frontier LLMs by up to 18.1 points, while simultaneously reducing token usage by 53.5% and agent steps by 43.7%. Results on SuperGPQA and TRQA confirm robustness across domains. Error analysis shows that reasoning failures and knowledge gaps co-occur in over 85% of cases, while diversity analysis reveals a clear dichotomy: retrieval tasks benefit from solution variety, whereas reasoning tasks favor consensus. Together, these findings demonstrate how implicit augmentation and structured refinement overcome the inefficiencies of explicit tool use and uniform aggregation. The code is available at https://github.com/ tangxiangru/Eigen-1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5aefb68e-2597-44aa-be9c-897fa0b1d9d3Builds on21
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat et al.ICML 2020 · 2,937 citations
Related papers
- HM-RAG: Hierarchical Multi-Agent Multimodal Retrieval Augmented GenerationPei Liu, Xin Liu, Ruoyu Yao, Junming Liu et al.ACM MM 2025 · 27 citations
- One Refiner to Unlock Them All: Inference-Time Reasoning Elicitation via Reinforcement Query RefinementYixiao Zhou, Dongzhou Cheng, Zhiliang Wu, Yi Yang et al.ACL 2026 · 3 citations
- BeamAggR: Beam Aggregation Reasoning over Multi-source Knowledge for Multi-hop Question AnsweringZheng Chu, Jingchang Chen, Qianglong Chen, Haotian Wang et al.ACL 2024 · 8 citations
- Iterative Multi-Granular RAG with Contextual Hierarchical GraphYanli Hu, Teng Liu, Zhuangyi Zhou, Weixin Zeng et al.AAAI 2026
- From Conflict to Consensus: Boosting Medical Reasoning via Multi-Round Agentic RAGWenhao Wu, Zhentao Tang, Yafu Li, Shixiong Kai et al.ICML 2026
