Optimistic Value Instructors for Cooperative Multi-Agent Reinforcement Learning
Chao Li, Yupeng Zhang, Jianqi Wang, Yujing Hu, Shaokang Dong, Wenbin Li, Tangjie Lv, Changjie Fan, Yang Gao
摘要
In cooperative multi-agent reinforcement learning, decentralized agents hold the promise of overcoming the combinatorial explosion of joint action space and enabling greater scalability. However, they are susceptible to a game-theoretic pathology called relative overgeneralization that shadows the optimal joint action. Although recent value-decomposition algorithms guide decentralized agents by learning a factored global action value function, the representational limitation and the inaccurate sampling of optimal joint actions during the learning process make this problem still. To address this limitation, this paper proposes a novel algorithm called Optimistic Value Instructors (OVI). The main idea behind OVI is to introduce multiple optimistic instructors into the valuedecomposition paradigm, which are capable of suggesting potentially optimal joint actions and rectifying the factored global action value function to recover these optimal actions. Specifically, the instructors maintain optimistic value estimations of per-agent local actions and thus eliminate the negative effects caused by other agents' exploratory or suboptimal non-cooperation, enabling accurate identification and suggestion of optimal joint actions. Based on the instructors' suggestions, the paper further presents two instructive constraints to rectify the factored global action value function to recover these optimal joint actions, thus overcoming the RO problem. Experimental evaluation of OVI on various cooperative multi-agent tasks demonstrates its superior performance against multiple baselines, highlighting its effectiveness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper3
- Weighted QMIX: Expanding Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement LearningTabish Rashid, Gregory Farquhar, Bei Peng, Shimon WhitesonNeurIPS 2020 · 被引用 1,960 次
- QPLEX: Duplex Dueling Multi-Agent Q-LearningJianhao Wang, Zhizhou Ren, Terry Liu, Yang Yu 等ICLR 2021 · 被引用 595 次
- UneVEn: Universal Value Exploration for Multi-Agent Reinforcement LearningTarun Gupta, Anuj Mahajan, Bei Peng, Wendelin Boehmer 等ICML 2021 · 被引用 59 次
相关 Paper
- Optimistic Multi-Agent Policy GradientWenshuai Zhao, Yi Zhao, Zhiyuan Li, Juho Kannala 等ICML 2024 · 被引用 7 次
- Potentially Optimal Joint Actions Recognition for Cooperative Multi-Agent Reinforcement LearningChang Huang, Shatong Zhu, Junqiao Zhao, Hongtu Zhou 等ICLR 2026
- Solving Homogeneous and Heterogeneous Cooperative Tasks with Greedy Sequential ExecutionShanqi Liu, Dong Xing, Pengjie Gu, Xinrun Wang 等ICLR 2024 · 被引用 2 次
- In-Context Fully Decentralized Cooperative Multi-Agent Reinforcement LearningChao Li, Bingkun Bao, Yang GaoNeurIPS 2025 · 被引用 2 次
- Locality Matters: A Scalable Value Decomposition Approach for Cooperative Multi-Agent Reinforcement LearningRoy Zohar, Shie Mannor, Guy TennenholtzAAAI 2022 · 被引用 11 次
