COLA: Consistent Learning with Opponent-Learning Awareness
Timon Willi, Alistair Letcher, Johannes Treutlein, Jakob N. Foerster
Abstract
Learning in general-sum games is unstable and frequently leads to socially undesirable (Pareto-dominated) outcomes. To mitigate this, Learning with Opponent-Learning Awareness (LOLA) introduced opponent shaping to this setting, by accounting for each agent's influence on their opponents' anticipated learning steps. However, the original LOLA formulation (and follow-up work) is inconsistent because LOLA models other agents as naive learners rather than LOLA agents. In previous work, this inconsistency was suggested as a cause of LOLA's failure to preserve stable fixed points (SFPs). First, we formalize consistency and show that higher-order LOLA (HOLA) solves LOLA's inconsistency problem if it converges. Second, we correct a claim made in the literature by Schäfer and Anandkumar (2019), proving that Competitive Gradient Descent (CGD) does not recover HOLA as a series expansion (and fails to solve the consistency problem). Third, we propose a new method called Consistent LOLA (COLA), which learns update functions that are consistent under mutual opponent shaping. It requires no more than second-order derivatives and learns consistent update functions even when HOLA fails to converge. However, we also prove that even consistent update functions do not preserve SFPs, contradicting the hypothesis that this shortcoming is caused by LOLA's inconsistency. Finally, in an empirical evaluation on a set of general-sum games, we find that COLA finds prosocial solutions and that it converges under a wider range of learning rates than HOLA and LOLA. We support the latter finding with a theoretical result for a simple game.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b4972b14-e2c8-4dc5-baba-a54e77321509Cited by top-tier papers22
- MADiff: Offline Multi-agent Learning with Diffusion ModelsZhengbang Zhu, Minghuan Liu, Liyuan Mao, Bingyi Kang et al.NeurIPS 2024 · 116 citations
- Model-Free Opponent ShapingChristopher Lu, Timon Willi, Christian A. Schröder de Witt, Jakob N. FoersterICML 2022 · 53 citations
- Proximal Learning With Opponent-Learning AwarenessStephen Zhao, Chris Lu, Roger B. Grosse, Jakob N. FoersterNeurIPS 2022 · 31 citations
- Adversarial Cheap TalkChris Lu, Timon Willi, Alistair Letcher, Jakob Nicolaus FoersterICML 2023 · 17 citations
- CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social DilemmasEmanuel Tewolde, Xiao Zhang, David Guzman Piedrahita, Vincent Conitzer et al.ICML 2026 · 15 citations
Builds on1
Related papers
- Advantage Alignment AlgorithmsJuan Agustin Duque, Milad Aghajohari, Tim Cooijmans, Razvan Ciuca et al.ICLR 2025
- LOQA: Learning with Opponent Q-Learning AwarenessMilad Aghajohari, Juan Agustin Duque, Tim Cooijmans, Aaron C. CourvilleICLR 2024 · 9 citations
- Competitive Gradient OptimizationAbhijeet Vyas, Brian Bullins, Kamyar AzizzadenesheliICML 2023 · 4 citations
- Reciprocal Reward Influence Encourages Cooperation From Self-Interested AgentsJohn L. Zhou, Weizhe Hong, Jonathan C. KaoNeurIPS 2024 · 5 citations
- Robust and Diverse Multi-Agent Learning via Rational Policy GradientNiklas Lauffer, Ameesh Shah, Micah Carroll, Sanjit A. Seshia et al.NeurIPS 2025 · 4 citations
