Coevolving with the Other You: Fine-Tuning LLM with Sequential Cooperative Multi-Agent Reinforcement Learning
Hao Ma, Tianyi Hu, Zhiqiang Pu, Boyin Liu, Xiaolin Ai, Yanyan Liang, Min Chen
摘要
Reinforcement learning (RL) has emerged as a pivotal technique for fine-tuning large language models (LLMs) on specific tasks. However, prevailing RL fine-tuning methods predominantly rely on PPO and its variants. Though these algorithms are effective in general RL settings, they often exhibit suboptimal performance and vulnerability to distribution collapse when applied to the fine-tuning of LLMs. In this paper, we propose CORY, extending the RL fine-tuning of LLMs to a sequential cooperative multi-agent reinforcement learning framework, to leverage the inherent coevolution and emergent capabilities of multi-agent systems. In CORY, the LLM to be fine-tuned is initially duplicated into two autonomous agents: a pioneer and an observer. The pioneer generates responses based on queries, while the observer generates responses using both the queries and the pioneer's responses. The two agents are trained together. During training, the agents exchange roles periodically, fostering cooperation and coevolution between them. Experiments evaluate CORY's performance by fine-tuning GPT-2 and Llama-2 under subjective and objective reward functions on the IMDB Review and GSM8K datasets, respectively. Results show that CORY outperforms PPO in terms of policy optimality, resistance to distribution collapse, and training robustness, thereby underscoring its potential as a superior methodology for refining LLMs in real-world applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task AutomationMengkang Hu, Yuhang Zhou, Wendong Fan, Yuzhou Nie 等NeurIPS 2025 · 被引用 158 次
- ReMA: Learning to Meta-Think for LLMs with Multi-agent Reinforcement LearningZiyu Wan, Yunxiang Li, Xiaoyu Wen, Yan Song 等NeurIPS 2025 · 被引用 76 次
- Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management for Large Language Model AgentsYi Yu, Liuyi Yao, Yuexiang Xie, Qingquan Tan 等ACL 2026 · 被引用 40 次
- Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language ModelsMickel Liu, Liwei Jiang, Yancheng Liang, Simon Du 等ICML 2026 · 被引用 34 次
- Heterogeneous Swarms: Jointly Optimizing Model Roles and Weights for Multi-LLM SystemsShangbin Feng, Zifeng Wang, Palash Goyal, Yike Wang 等NeurIPS 2025 · 被引用 26 次
它引用的顶会 Paper9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch 等ICML 2023 · 被引用 2,601 次
- Pre-Trained Language Models for Interactive Decision-MakingShuang Li, Xavier Puig, Chris Paxton, Yilun Du 等NeurIPS 2022 · 被引用 341 次
相关 Paper
- LLM Collaboration with Multi-Agent Reinforcement LearningShuo Liu, Zeyu Liang, Xueguang Lyu, Christopher AmatoAAAI 2026
- Multi-Agent Collaboration via Evolving OrchestrationYufan Dang, Chen Qian, Xueheng Luo, Jingru Fan 等NeurIPS 2025 · 被引用 118 次
- CoPE: A Framework for Optimizing Coordination between Planning and Execution in LLM-based AgentsHuanxi Liu, Kun Hu, Qiang Wang, Yuanzhao Zhai 等ICML 2026
- Multiagent Finetuning: Self Improvement with Diverse Reasoning ChainsVighnesh Subramaniam, Yilun Du, Joshua B. Tenenbaum, Antonio Torralba 等ICLR 2025
- Online Preference Alignment for Language Models via Count-based ExplorationChenjia Bai, Yang Zhang, Shuang Qiu, Qiaosheng Zhang 等ICLR 2025
