DECOR: Learning to Decompose and Collaborate in Deep Search via Multi-Agent Reinforcement Learning
Ruiqing Chen, Zekun Zhang, Gongduo Zhang, Lihong Gu, Lin Zhou
Abstract
Monolithic agents in deep search often suffer from "cognitive overload," while existing multi-agent approaches mostly rely on frozen models that cannot learn from collaboration failures. To bridge this gap, we propose (compose and llaborate via ole-specialized agents), a framework formulating deep search as a Multi-Agent Reinforcement Learning (MARL) problem. DECOR functionally decomposes the task into three specialized roles: a to navigate, a to curate a noise-reduced memory, and an for synthesis. Unlike training-free orchestration, we jointly optimize these agents using a hybrid reward strategy that harmonizes role-specific intrinsic feedback with team-level outcome signals. Experiments on seven benchmarks show that DECOR significantly outperforms strong monolithic baselines, demonstrating the necessity of learning-based functional decomposition in handling cognitive overload.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on8
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric MemoriesAlex Mallen, Akari Asai, Victor Zhong, Rajarshi Das et al.ACL 2023 · 233 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
Related papers
- MASER: Multi-Agent Reinforcement Learning with Subgoals Generated from Experience Replay BufferJeewon Jeon, Woojun Kim, Whiyoung Jung, Youngchul SungICML 2022 · 53 citations
- Expert-Inspired Multi-Agent Coordination for Multi-Objective Molecular OptimizationDaojian Zeng, Tianle Li, Jiahao Yang, Jiacai Yi et al.AAAI 2026
- DeCOM: Decomposed Policy for Constrained Cooperative Multi-Agent Reinforcement LearningZhaoxing Yang, Haiming Jin, Rong Ding, Haoyi You et al.AAAI 2023 · 5 citations
- HAVEN: Hierarchical Cooperative Multi-Agent Reinforcement Learning with Dual Coordination MechanismZhiwei Xu, Yunpeng Bai, Bin Zhang, Dapeng Li et al.AAAI 2023 · 46 citations
- MARS²: Scaling Multi-Agent Tree Search via Reinforcement Learning for Code GenerationPengfei Li, Shijie Wang, Fangyuan Li, Yikun Fu et al.ACL 2026
