TeamTR: Trust-Region Fine-Tuning for Multi-Agent LLM Coordination
Yi Xie, Siao Liu, Falong FAN, Yuanqi Yao, Siyang Cao, Yue Zhao, Bo Liu
摘要
Multi-agent LLM systems can improve reasoning and tool use, yet recent evidence shows their gains are often unstable and sensitive to interaction design. A promising direction is to train collaboration, but team post-training introduces a moving-target effect: when agents interact through a shared context, updating one agent shifts the context distribution faced by the others, which can regress coordination under naive sequential updates. We propose TeamTR, a trust-region framework for fine-tuning heterogeneous LLM teams that explicitly controls this occupancy shift. TeamTR evaluates each agent update on rollouts from the intermediate team induced by partially applied updates, and enforces per-agent trust regions via a token-decomposed reverse KL that is directly monitorable from those rollouts. This yields population-level per-update and per-stage improvement lower bounds whose functional form applies to any realized update order, and motivates a practical certificate proxy computed from logged surrogates and KL terms. We instantiate TeamTR for router-based text handoff with sequence-level returns and bounded group-normalized advantages, and show empirically that it mitigates coordination regressions, improves training stability across heterogeneous teams, and supports modular component replacement via a trust-region alignment step.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper22
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model SocietyGuohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin 等NeurIPS 2023 · 被引用 1,975 次
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris 等UIST 2023 · 被引用 1,882 次
相关 Paper
- Trust Region Policy Optimisation in Multi-Agent Reinforcement LearningJakub Grudzien Kuba, Ruiqing Chen, Muning Wen, Ying Wen 等ICLR 2022 · 被引用 367 次
- Trust-Region Adaptive Policy OptimizationMingyu Su, Jian Guan, Yuxian Gu, Minlie Huang 等ICLR 2026 · 被引用 2 次
- TROLL: Trust Regions Improve Reinforcement Learning for Large Language ModelsPhilipp Becker, Niklas Freymuth, Serge Thilges, Fabian Otto 等ICLR 2026 · 被引用 8 次
- AlphaRouter: Token-level Routing Between SLM and LLM with Reinforcement Learning and Tree SearchSiteng Liao, Yuzhu Liang, Hengzhong Rao, Xizhao Luo 等ICML 2026
- RobustRL: Role-Based Fault Tolerance System for RL Post-TrainingZhenqian Chen, Baoquan Zhong, Xiang Li, Qing Dai 等OSDI 2026
