NonZero: Interaction-Guided Exploration for Multi-Agent Monte Carlo Tree Search
Sizhe Tang, Zuyuan Zhang, Mahdi Imani, Tian Lan
摘要
Monte Carlo Tree Search (MCTS) scales poorly in cooperative multi-agent domains because expansion must consider an exponentially large set of joint actions, severely limiting exploration under realistic search budgets. We propose NonZero, which keeps multi-agent MCTS tractable by running surrogate-guided selection over a low-dimensional nonlinear representation using an interaction-guided proposal rule, instead of directly exploring the full joint-action space. Our exploration uses an interaction score: single-agent deviations are ranked by predicted gain, while two-agent deviations are scored by a mixed-difference measure that reveals coordination benefits even when no single agent can improve alone. We formalize candidate proposal as a bandit problem over local deviations and derive a proposal rule, NonUCT, with a sublinear local-regret guarantee for reaching approximate graph-local optima without enumerating the joint-action space. Empirically, NonZero improves sample efficiency and final performance on MatGame, SMAC, and SMACv2 relative to strong model-based and model-free baselines under matched search budgets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Weighted QMIX: Expanding Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement LearningTabish Rashid, Gregory Farquhar, Bei Peng, Shimon WhitesonNeurIPS 2020 · 被引用 1,960 次
- Mastering Atari Games with Limited DataWeirui Ye, Shaohuai Liu, Thanard Kurutach, Pieter Abbeel 等NeurIPS 2021 · 被引用 345 次
- DOP: Off-Policy Multi-Agent Decomposed Policy GradientsYihan Wang, Beining Han, Tonghan Wang, Heng Dong 等ICLR 2021 · 被引用 208 次
- Learning Nearly Decomposable Value Functions Via Communication MinimizationTonghan Wang, Jianhao Wang, Chongyi Zheng, Chongjie ZhangICLR 2020 · 被引用 170 次
- Learning and Planning in Complex Action SpacesThomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Mohammadamin Barekatain 等ICML 2021 · 被引用 99 次
相关 Paper
- MALinZero: Efficient Low-Dimensional Search for Mastering Complex Multi-Agent PlanningSizhe Tang, Jiayu Chen, Tian LanNeurIPS 2025 · 被引用 9 次
- Efficient Multi-agent Reinforcement Learning by PlanningQihan Liu, Jianing Ye, Xiaoteng Ma, Jun Yang 等ICLR 2024 · 被引用 18 次
- Learning Joint Behaviors with Large VariationsTianxu Li, Kun ZhuAAAI 2025 · 被引用 2 次
- Non-Linear Coordination GraphsYipeng Kang, Tonghan Wang, Qianlan Yang, Xiaoran Wu 等NeurIPS 2022 · 被引用 14 次
- Scalable Safe Policy Improvement for Factored Multi-Agent MDPsFederico Bianchi, Edoardo Zorzi, Alberto Castellini, Thiago D. Simão 等ICML 2024 · 被引用 3 次
