Learning to Cooperate with Minimal Observability
Chin-wing Leung, Paolo Turrini, Fernando P. Santos, Mirco Musolesi
摘要
Cooperation among independent learning agents is desirable as it enables reaching collectively rewarding states. Recent work has shown that artificial agents can learn to act pro-socially without the need for predefined cooperative preferences or behavioural heuristics, provided that they can observe others' actions or policies and select them as partners accordingly. This paper relaxes this constraint, studying reinforcement learning (RL) agents operating with only minimal information about others' behaviour. We propose a novel `Observer Model', where agents gain insights from direct experience and limited, indirect observations. We show that direct experience alone cannot sustain cooperation, particularly in large societies. However, even minimal observations of third-party interactions, allowing as few as one observer per gameplay, lead to significant improvements, enabling the population to achieve and sustain robust cooperation across varying population sizes. Through numerical analysis, we show the co-evolution of strategy and interaction structure and disentangle how learning happens under various settings. Analysing the partner selection graph, we identify the reasons for cooperation to emerge, and we explore how different learning and exploration rates affect the outcome of social dilemmas played among RL agents.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper1
相关 Paper
- Reciprocal Reward Influence Encourages Cooperation From Self-Interested AgentsJohn L. Zhou, Weizhe Hong, Jonathan C. KaoNeurIPS 2024 · 被引用 5 次
- Learning Altruistic Behaviours in Reinforcement Learning without External RewardsTim Franzmeyer, Mateusz Malinowski, João F. HenriquesICLR 2022 · 被引用 10 次
- Emergence of Punishment in Social Dilemma with Environmental FeedbackZhen Wang, Zhao Song, Chen Shen, Shuyue HuAAAI 2023 · 被引用 41 次
- Multi-agent cooperation through learning-aware policy gradientsAlexander Meulemans, Seijin Kobayashi, Johannes von Oswald, Nino Scherrer 等ICLR 2025
- Learning to Communicate Implicitly by ActionsZheng Tian, Shihao Zou, Ian Davies, Tim Warr 等AAAI 2020 · 被引用 35 次
