LOQA: Learning with Opponent Q-Learning Awareness
Milad Aghajohari, Juan Agustin Duque, Tim Cooijmans, Aaron C. Courville
摘要
In various real-world scenarios, interactions among agents often resemble the dynamics of general-sum games, where each agent strives to optimize its own utility. Despite the ubiquitous relevance of such settings, decentralized machine learning algorithms have struggled to find equilibria that maximize individual utility while preserving social welfare. In this paper we introduce Learning with Opponent Q-Learning Awareness (LOQA), a novel, decentralized reinforcement learning algorithm tailored to optimizing an agent's individual utility while fostering cooperation among adversaries in partially competitive environments. LOQA assumes the opponent samples actions proportionally to their action-value function Q. Experimental results demonstrate the effectiveness of LOQA at achieving state-of-the-art performance in benchmark scenarios such as the Iterated Prisoner's Dilemma and the Coin Game. LOQA achieves these outcomes with a significantly reduced computational footprint, making it a promising approach for practical multi-agent applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Multi-agent cooperation through learning-aware policy gradientsAlexander Meulemans, Seijin Kobayashi, Johannes von Oswald, Nino Scherrer 等ICLR 2025
- Advantage Alignment AlgorithmsJuan Agustin Duque, Milad Aghajohari, Tim Cooijmans, Razvan Ciuca 等ICLR 2025
- Towards Sustainable Investment Policies Informed by Opponent ShapingJuan Agustin Duque, Razvan Ciuca, Ayoub Echchahed, Hugo Larochelle 等ICLR 2026
它引用的顶会 Paper4
- A Policy Gradient Algorithm for Learning to Learn in Multiagent Reinforcement LearningDong-Ki Kim, Miao Liu, Matthew Riemer, Chuangchuang Sun 等ICML 2021 · 被引用 66 次
- COLA: Consistent Learning with Opponent-Learning AwarenessTimon Willi, Alistair Letcher, Johannes Treutlein, Jakob N. FoersterICML 2022 · 被引用 61 次
- Model-Free Opponent ShapingChristopher Lu, Timon Willi, Christian A. Schröder de Witt, Jakob N. FoersterICML 2022 · 被引用 53 次
- Proximal Learning With Opponent-Learning AwarenessStephen Zhao, Chris Lu, Roger B. Grosse, Jakob N. FoersterNeurIPS 2022 · 被引用 31 次
相关 Paper
- Locality Matters: A Scalable Value Decomposition Approach for Cooperative Multi-Agent Reinforcement LearningRoy Zohar, Shie Mannor, Guy TennenholtzAAAI 2022 · 被引用 11 次
- Learning to Incentivize Other Learning AgentsJiachen Yang, Ang Li, Mehrdad Farajtabar, Peter Sunehag 等NeurIPS 2020 · 被引用 105 次
- Decentralized Q-learning in Zero-sum Markov GamesMuhammed O. Sayin, Kaiqing Zhang, David S. Leslie, Tamer Basar 等NeurIPS 2021 · 被引用 105 次
- Reciprocal Reward Influence Encourages Cooperation From Self-Interested AgentsJohn L. Zhou, Weizhe Hong, Jonathan C. KaoNeurIPS 2024 · 被引用 5 次
- Weighted QMIX: Expanding Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement LearningTabish Rashid, Gregory Farquhar, Bei Peng, Shimon WhitesonNeurIPS 2020 · 被引用 1,960 次
