Proximal Learning With Opponent-Learning Awareness
Stephen Zhao, Chris Lu, Roger B. Grosse, Jakob N. Foerster
Abstract
Learning With Opponent-Learning Awareness (LOLA) (Foerster et al. [2018a]) is a multi-agent reinforcement learning algorithm that typically learns reciprocity-based cooperation in partially competitive environments. However, LOLA often fails to learn such behaviour on more complex policy spaces parameterized by neural networks, partly because the update rule is sensitive to the policy parameterization. This problem is especially pronounced in the opponent modeling setting, where the opponent's policy is unknown and must be inferred from observations; in such settings, LOLA is ill-specified because behaviourally equivalent opponent policies can result in non-equivalent updates. To address this shortcoming, we reinterpret LOLA as approximating a proximal operator, and then derive a new algorithm, proximal LOLA (POLA), which uses the proximal formulation directly. Unlike LOLA, the POLA updates are parameterization invariant, in the sense that when the proximal objective has a unique optimum, behaviourally equivalent policies result in behaviourally equivalent updates. We then present practical approximations to the ideal POLA update, which we evaluate in several partially competitive environments with function approximation and opponent modeling. This empirically demonstrates that POLA achieves reciprocity-based cooperation more reliably than LOLA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9816c1e0-6774-4843-8497-e147849d1cd3Cited by top-tier papers11
- MADiff: Offline Multi-agent Learning with Diffusion ModelsZhengbang Zhu, Minghuan Liu, Liyuan Mao, Bingyi Kang et al.NeurIPS 2024 · 116 citations
- LOQA: Learning with Opponent Q-Learning AwarenessMilad Aghajohari, Juan Agustin Duque, Tim Cooijmans, Aaron C. CourvilleICLR 2024 · 9 citations
- Opponent Modeling with In-context SearchYuheng Jing, Bingyun Liu, Kai Li, Yifan Zang et al.NeurIPS 2024 · 8 citations
- Towards Offline Opponent Modeling with In-context LearningYuheng Jing, Kai Li, Bingyun Liu, Yifan Zang et al.ICLR 2024 · 7 citations
- Efficient Adaptation in Mixed-Motive Environments via Hierarchical Opponent Modeling and PlanningYizhe Huang, Anji Liu, Fanqi Kong, Yaodong Yang et al.ICML 2024 · 5 citations
Builds on4
- COLA: Consistent Learning with Opponent-Learning AwarenessTimon Willi, Alistair Letcher, Johannes Treutlein, Jakob N. FoersterICML 2022 · 61 citations
- Model-Free Opponent ShapingChristopher Lu, Timon Willi, Christian A. Schröder de Witt, Jakob N. FoersterICML 2022 · 53 citations
- Mirror Learning: A Unifying Framework of Policy OptimisationJakub Grudzien Kuba, Christian A. Schröder de Witt, Jakob N. FoersterICML 2022 · 41 citations
- Minimax Optimization with Smooth Algorithmic AdversariesTanner Fiez, Chi Jin, Praneeth Netrapalli, Lillian J. RatliffICLR 2022 · 11 citations
Related papers
- Adversarial Policy Training against Deep Reinforcement LearningXian Wu, Wenbo Guo, Hua Wei, Xinyu XingUSENIX Security 2021 · 19 citations
- Opponent Modeling based on Subgoal InferenceXiaopeng Yu, Jiechuan Jiang, Zongqing LuNeurIPS 2024 · 7 citations
- Reciprocal Reward Influence Encourages Cooperation From Self-Interested AgentsJohn L. Zhou, Weizhe Hong, Jonathan C. KaoNeurIPS 2024 · 5 citations
- Wasserstein Flow Meets Replicator Dynamics: A Mean-Field Analysis of Representation Learning in Actor-CriticYufeng Zhang, Siyu Chen, Zhuoran Yang, Michael I. Jordan et al.NeurIPS 2021 · 6 citations
- Proximal Distilled Evolutionary Reinforcement LearningCristian Bodnar, Ben Day, Pietro LióAAAI 2020 · 101 citations
