Self-Play Q-Learners Can Provably Collude in the Iterated Prisoner's Dilemma
Quentin Bertrand, Juan Agustin Duque, Emilio Calvano, Gauthier Gidel
Abstract
A growing body of computational studies shows that simple machine learning agents converge to cooperative behaviors in social dilemmas, such as collusive price-setting in oligopoly markets, raising questions about what drives this outcome. In this work, we provide theoretical foundations for this phenomenon in the context of self-play multi-agent Q-learners in the iterated prisoner's dilemma. We characterize broad conditions under which such agents provably learn the cooperative Pavlov (win-stay, lose-shift) policy rather than the Pareto-dominated "always defect" policy. We validate our theoretical results through additional experiments, demonstrating their robustness across a broader class of deep learning algorithms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 41257edb-b841-4801-bc7d-036472ccf0ecCited by top-tier papers2
- CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social DilemmasEmanuel Tewolde, Xiao Zhang, David Guzman Piedrahita, Vincent Conitzer et al.ICML 2026 · 15 citations
- Convex Markov Games: A New Frontier for Multi-Agent Reinforcement LearningIan Gemp, Andreas Alexander Haupt, Luke Marris, Siqi Liu et al.ICML 2025
Builds on1
Related papers
- Reciprocal Reward Influence Encourages Cooperation From Self-Interested AgentsJohn L. Zhou, Weizhe Hong, Jonathan C. KaoNeurIPS 2024 · 5 citations
- Stability of Multi-Agent Learning in Competitive Networks: Delaying the Onset of ChaosAamal Abbas Hussain, Francesco BelardinelliAAAI 2024 · 4 citations
- Similarity-based cooperative equilibriumCaspar Oesterheld, Johannes Treutlein, Roger B. Grosse, Vincent Conitzer et al.NeurIPS 2023 · 8 citations
- Emergent Fast-Slow Dynamics in Multi-Agent Q-Learning for Networked Stochastic GamesYuxin Geng, Wolfram Barfuss, Xingru ChenAAAI 2026
- Exploration-Exploitation in Multi-Agent Competition: Convergence with Bounded RationalityStefanos Leonardos, Georgios Piliouras, Kelly SpendloveNeurIPS 2021 · 43 citations
