OPtions as REsponses: Grounding behavioural hierarchies in multi-agent reinforcement learning
Alexander Vezhnevets, Yuhuai Wu, Maria K. Eckstein, Rémi Leblond, Joel Z. Leibo
摘要
This paper investigates generalisation in multiagent games, where the generality of the agent can be evaluated by playing against opponents it hasn't seen during training. We propose two new games with concealed information and complex, non-transitive reward structure (think rock/paper/scissors). It turns out that most current deep reinforcement learning methods fail to efficiently explore the strategy space, thus learning policies that generalise poorly to unseen opponents. We then propose a novel hierarchical agent architecture, where the hierarchy is grounded in the game-theoretic structure of the game -the top level chooses strategic responses to opponents, while the low level implements them into policy over primitive actions. This grounding facilitates credit assignment across the levels of hierarchy. Our experiments show that the proposed hierarchical agent is capable of generalisation to unseen opponents, while conventional baselines fail to generalise whatsoever.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Scalable Evaluation of Multi-Agent Reinforcement Learning with Melting PotJoel Z. Leibo, Edgar A. Duéñez-Guzmán, Alexander Vezhnevets, John P. Agapiou 等ICML 2021 · 被引用 134 次
- N-agent Ad Hoc TeamworkCaroline Wang, Arrasy Rahman, Ishan Durugkar, Elad Liebman 等NeurIPS 2024 · 被引用 21 次
- NeuPL: Neural Population LearningSiqi Liu, Luke Marris, Daniel Hennes, Josh Merel 等ICLR 2022 · 被引用 19 次
- Simplex Neural Population Learning: Any-Mixture Bayes-Optimality in Symmetric Zero-sum GamesSiqi Liu, Marc Lanctot, Luke Marris, Nicolas HeessICML 2022 · 被引用 12 次
- SocialJax: An Evaluation Suite for Multi-agent Reinforcement Learning in Sequential Social DilemmasZihao Guo, Shuqing Shi, Richard Willis, Tristan Tomilin 等ICLR 2026 · 被引用 11 次
它引用的顶会 Paper1
相关 Paper
- Multi-Agent Actor-Critic with Hierarchical Graph Attention NetworkHeechang Ryu, Hayong Shin, Jinkyoo ParkAAAI 2020 · 被引用 143 次
- Accelerating Task Generalisation with Multi-Level Skill HierarchiesThomas P. Cannon, Özgür SimsekICLR 2025
- Deep Reinforcement Learning for General Game PlayingAdrian Goldwaser, Michael ThielscherAAAI 2020 · 被引用 46 次
- Environmental drivers of systematicity and generalization in a situated agentFelix Hill, Andrew K. Lampinen, Rosalia Schneider, Stephen Clark 等ICLR 2020 · 被引用 109 次
- Deep Coordination GraphsWendelin Boehmer, Vitaly Kurin, Shimon WhitesonICML 2020 · 被引用 209 次
