CoEvol: Constructing Better Responses for Instruction Finetuning through Multi-Agent Cooperation
Renhao Li, Minghuan Tan, Derek F. Wong, Min Yang
摘要
In recent years, instruction fine-tuning (IFT) on large language models (LLMs) has garnered considerable attention to enhance model performance on unseen tasks. Attempts have been made on automatic construction and effective selection for IFT data. However, we posit that previous methods have not fully harnessed the potential of LLMs for enhancing data quality. The responses within IFT data could be further enhanced by leveraging the capabilities of LLMs themselves. In this paper, we propose COEVOL, an LLM-based multiagent cooperation framework for the improvement of responses for instructions. To effectively refine the responses, we develop an iterative framework following a debate-adviseedit-judge paradigm. A two-stage multi-agent debate strategy is further devised to ensure the diversity and reliability of editing suggestions within the framework. Empirically, models equipped with COEVOL outperform competitive baselines evaluated by MT-Bench and Al-pacaEval, demonstrating its effectiveness in enhancing instruction-following capabilities for LLMs. 1 * Equal contribution. † Under the Joint Ph.D. Program between UM and SIAT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human EvaluationJiaju Chen, Yuxuan Lu, Xiaojie Wang, Huimin Zeng 等ACL 2026 · 被引用 30 次
- Collaborative Reasoner: Self-Improving Social Agents with Synthetic ConversationsAnsong Ni, Ruta Desai, Yang Li, Xinjie Lei 等NeurIPS 2025 · 被引用 7 次
- BELLE: A Bi-Level Multi-Agent Reasoning Framework for Multi-Hop Question AnsweringTaolin Zhang, Dongyang Li, Qizhou Chen, Chengyu Wang 等ACL 2025
它引用的顶会 Paper22
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris 等UIST 2023 · 被引用 1,882 次
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum 等ICML 2024 · 被引用 1,562 次
- LIMA: Less Is More for AlignmentChunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer 等NeurIPS 2023 · 被引用 1,486 次
相关 Paper
- Evoke: Evoking Critical Thinking Abilities in LLMs via Reviewer-Author Prompt EditingXinyu Hu, Pengfei Tang, Simiao Zuo, Zihan Wang 等ICLR 2024 · 被引用 14 次
- Automatic Instruction Evolving for Large Language ModelsWeihao Zeng, Can Xu, Yingxiu Zhao, Jian-Guang Lou 等EMNLP 2024 · 被引用 2 次
- MMIFEvol: Towards Evolutionary Multimodal Instruction FollowingHaoyu Wang, Sihang Jiang, Xiangru Zhu, Yuyan Chen 等AAAI 2026 · 被引用 1 次
- Coevolving with the Other You: Fine-Tuning LLM with Sequential Cooperative Multi-Agent Reinforcement LearningHao Ma, Tianyi Hu, Zhiqiang Pu, Boyin Liu 等NeurIPS 2024 · 被引用 54 次
- From Selection to Refinement: Iterative Optimization for Instruction DataHang Hu, Ziyan Liu, Rujie Wen, Ruihui Hou 等ACL 2026
