Model-Free Opponent Shaping
Christopher Lu, Timon Willi, Christian A. Schröder de Witt, Jakob N. Foerster
摘要
In general-sum games, the interaction of selfinterested learning agents commonly leads to collectively worst-case outcomes, such as defectdefect in the iterated prisoner's dilemma (IPD). To overcome this, some methods, such as Learning with Opponent-Learning Awareness (LOLA), shape their opponents' learning process. However, these methods are myopic since only a small number of steps can be anticipated, are asymmetric since they treat other agents as naive learners, and require the use of higher-order derivatives, which are calculated through white-box access to an opponent's differentiable learning algorithm. To address these issues, we propose Model-Free Opponent Shaping (M-FOS). M-FOS learns in a meta-game in which each meta-step is an episode of the underlying ("inner") game. The meta-state consists of the inner policies, and the meta-policy produces a new inner policy to be used in the next episode. M-FOS then uses generic model-free optimisation methods to learn meta-policies that accomplish long-horizon opponent shaping. Empirically, M-FOS near-optimally exploits naive learners and other, more sophisticated algorithms from the literature. For example, to the best of our knowledge, it is the first method to learn the well-known Zero-Determinant (ZD) extortion strategy in the IPD. In the same settings, M-FOS leads to socially optimal outcomes under meta-self-play. Finally, we show that M-FOS can be scaled to highdimensional settings. Project code is available at: https://github.com/luchris429/ Model-Free-Opponent-Shaping .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper29
- COLA: Consistent Learning with Opponent-Learning AwarenessTimon Willi, Alistair Letcher, Johannes Treutlein, Jakob N. FoersterICML 2022 · 被引用 61 次
- Proximal Learning With Opponent-Learning AwarenessStephen Zhao, Chris Lu, Roger B. Grosse, Jakob N. FoersterNeurIPS 2022 · 被引用 31 次
- Influencing Long-Term Behavior in Multiagent Reinforcement LearningDong-Ki Kim, Matthew Riemer, Miao Liu, Jakob N. Foerster 等NeurIPS 2022 · 被引用 29 次
- Oracles & Followers: Stackelberg Equilibria in Deep Multi-Agent Reinforcement LearningMatthias Gerstgrasser, David C. ParkesICML 2023 · 被引用 27 次
- Discovering Temporally-Aware Reinforcement Learning AlgorithmsMatthew Thomas Jackson, Chris Lu, Louis Kirsch, Robert Tjarko Lange 等ICLR 2024 · 被引用 23 次
它引用的顶会 Paper3
- A Policy Gradient Algorithm for Learning to Learn in Multiagent Reinforcement LearningDong-Ki Kim, Miao Liu, Matthew Riemer, Chuangchuang Sun 等ICML 2021 · 被引用 66 次
- COLA: Consistent Learning with Opponent-Learning AwarenessTimon Willi, Alistair Letcher, Johannes Treutlein, Jakob N. FoersterICML 2022 · 被引用 61 次
- Model-Based Opponent ModelingXiaopeng Yu, Jiechuan Jiang, Wanpeng Zhang, Haobin Jiang 等NeurIPS 2022 · 被引用 56 次
相关 Paper
- Reciprocal Reward Influence Encourages Cooperation From Self-Interested AgentsJohn L. Zhou, Weizhe Hong, Jonathan C. KaoNeurIPS 2024 · 被引用 5 次
- Advantage Alignment AlgorithmsJuan Agustin Duque, Milad Aghajohari, Tim Cooijmans, Razvan Ciuca 等ICLR 2025
- Opponent Shaping in LLM AgentsMarta Emili Garcia Segura, Stephen Hailes, Mirco MusolesiICLR 2026 · 被引用 3 次
- Robust and Diverse Multi-Agent Learning via Rational Policy GradientNiklas Lauffer, Ameesh Shah, Micah Carroll, Sanjit A. Seshia 等NeurIPS 2025 · 被引用 4 次
- Multi-agent cooperation through learning-aware policy gradientsAlexander Meulemans, Seijin Kobayashi, Johannes von Oswald, Nino Scherrer 等ICLR 2025
