Learning to Cooperate with Minimal Observability
Chin-wing Leung, Paolo Turrini, Fernando P. Santos, Mirco Musolesi
Abstract
Cooperation among independent learning agents is desirable as it enables reaching collectively rewarding states. Recent work has shown that artificial agents can learn to act pro-socially without the need for predefined cooperative preferences or behavioural heuristics, provided that they can observe others' actions or policies and select them as partners accordingly. This paper relaxes this constraint, studying reinforcement learning (RL) agents operating with only minimal information about others' behaviour. We propose a novel `Observer Model', where agents gain insights from direct experience and limited, indirect observations. We show that direct experience alone cannot sustain cooperation, particularly in large societies. However, even minimal observations of third-party interactions, allowing as few as one observer per gameplay, lead to significant improvements, enabling the population to achieve and sustain robust cooperation across varying population sizes. Through numerical analysis, we show the co-evolution of strategy and interaction structure and disentangle how learning happens under various settings. Analysing the partner selection graph, we identify the reasons for cooperation to emerge, and we explore how different learning and exploration rates affect the outcome of social dilemmas played among RL agents.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on1
Related papers
- Reciprocal Reward Influence Encourages Cooperation From Self-Interested AgentsJohn L. Zhou, Weizhe Hong, Jonathan C. KaoNeurIPS 2024 · 5 citations
- Learning Altruistic Behaviours in Reinforcement Learning without External RewardsTim Franzmeyer, Mateusz Malinowski, João F. HenriquesICLR 2022 · 10 citations
- Emergence of Punishment in Social Dilemma with Environmental FeedbackZhen Wang, Zhao Song, Chen Shen, Shuyue HuAAAI 2023 · 41 citations
- Multi-agent cooperation through learning-aware policy gradientsAlexander Meulemans, Seijin Kobayashi, Johannes von Oswald, Nino Scherrer et al.ICLR 2025
- Learning to Communicate Implicitly by ActionsZheng Tian, Shihao Zou, Ian Davies, Tim Warr et al.AAAI 2020 · 35 citations
