UneVEn: Universal Value Exploration for Multi-Agent Reinforcement Learning
Tarun Gupta, Anuj Mahajan, Bei Peng, Wendelin Boehmer, Shimon Whiteson
摘要
This paper focuses on cooperative value-based multi-agent reinforcement learning (MARL) in the paradigm of centralized training with decentralized execution (CTDE). Current state-of-the-art value-based MARL methods leverage CTDE to learn a centralized joint-action value function as a monotonic mixing of each agent's utility function, which enables easy decentralization. However, this monotonic restriction leads to inefficient exploration in tasks with nonmonotonic returns due to suboptimal approximations of the values of joint actions. To address this, we present a novel MARL approach called Universal Value Exploration (UneVEn), which uses universal successor features (USFs) to learn policies of tasks related to the target task, but with simpler reward functions in a sample efficient manner. UneVEn uses novel action-selection schemes between randomly sampled related tasks during exploration, which enables the monotonic joint-action value function of the target task to place more importance on useful joint actions. Empirical results on a challenging cooperative predator-prey task requiring significant coordination amongst agents show that UneVEn significantly outperforms state-of-the-art baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- LDSA: Learning Dynamic Subtask Assignment in Cooperative Multi-Agent Reinforcement LearningMingyu Yang, Jian Zhao, Xunhan Hu, Wengang Zhou 等NeurIPS 2022 · 被引用 61 次
- Regularized Softmax Deep Multi-Agent Q-LearningLing Pan, Tabish Rashid, Bei Peng, Longbo Huang 等NeurIPS 2021 · 被引用 50 次
- PAC: Assisted Value Factorization with Counterfactual Predictions in Multi-Agent Reinforcement LearningHanhan Zhou, Tian Lan, Vaneet AggarwalNeurIPS 2022 · 被引用 47 次
- Tesseract: Tensorised Actors for Multi-Agent Reinforcement LearningAnuj Mahajan, Mikayel Samvelyan, Lei Mao, Viktor Makoviychuk 等ICML 2021 · 被引用 38 次
- Automatic Grouping for Efficient Cooperative Multi-Agent Reinforcement LearningYifan Zang, Jinmin He, Kai Li, Haobo Fu 等NeurIPS 2023 · 被引用 37 次
它引用的顶会 Paper3
- Deep Coordination GraphsWendelin Boehmer, Vitaly Kurin, Shimon WhitesonICML 2020 · 被引用 209 次
- Fast Task Inference with Variational Intrinsic Successor FeaturesSteven Hansen, Will Dabney, André Barreto, David Warde-Farley 等ICLR 2020 · 被引用 176 次
- Influence-Based Multi-Agent ExplorationTonghan Wang, Jianhao Wang, Yi Wu, Chongjie ZhangICLR 2020 · 被引用 156 次
相关 Paper
- Multi-Agent Guided Policy OptimizationYueheng Li, Guangming Xie, Zongqing LuICLR 2026 · 被引用 4 次
- Discovering Generalizable Multi-agent Coordination Skills from Multi-task Offline DataFuxiang Zhang, Chengxing Jia, Yi-Chen Li, Lei Yuan 等ICLR 2023
- Beyond Monotonicity: Revisiting Factorization Principles in Multi-Agent Q-LearningTianmeng Hu, Yongzheng Cui, Rui Tang, Biao Luo 等AAAI 2026
- Retaining Suboptimal Actions to Follow Shifting Optima in Multi-Agent Reinforcement LearningYonghyeon Jo, Sunwoo Lee, Seungyul HanICLR 2026 · 被引用 5 次
- Solving Homogeneous and Heterogeneous Cooperative Tasks with Greedy Sequential ExecutionShanqi Liu, Dong Xing, Pengjie Gu, Xinrun Wang 等ICLR 2024 · 被引用 2 次
