Learning to Incentivize Other Learning Agents
Jiachen Yang, Ang Li, Mehrdad Farajtabar, Peter Sunehag, Edward Hughes, Hongyuan Zha
摘要
The challenge of developing powerful and general Reinforcement Learning (RL) agents has received increasing attention in recent years. Much of this effort has focused on the single-agent setting, in which an agent maximizes a predefined extrinsic reward function. However, a long-term question inevitably arises: how will such independent agents cooperate when they are continually learning and acting in a shared multi-agent environment? Observing that humans often provide incentives to influence others' behavior, we propose to equip each RL agent in a multi-agent environment with the ability to give rewards directly to other agents, using a learned incentive function. Each agent learns its own incentive function by explicitly accounting for its impact on the learning of recipients and, through them, the impact on its own extrinsic objective. We demonstrate in experiments that such agents significantly outperform standard RL and opponent-shaping agents in challenging general-sum Markov games, often by finding a near-optimal division of labor. Our work points toward more opportunities and challenges along the path to ensure the common good in a multi-agent future.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Human-Timescale Adaptation in an Open-Ended Task SpaceJakob Bauer, Kate Baumli, Feryal M. P. Behbahani, Avishkar Bhoopchand 等ICML 2023 · 被引用 155 次
- Influencing Long-Term Behavior in Multiagent Reinforcement LearningDong-Ki Kim, Matthew Riemer, Miao Liu, Jakob N. Foerster 等NeurIPS 2022 · 被引用 29 次
- Information Design in Multi-Agent Reinforcement LearningYue Lin, Wenhao Li, Hongyuan Zha, Baoxiang WangNeurIPS 2023 · 被引用 25 次
- Contextual Bilevel Reinforcement Learning for Incentive AlignmentVinzenz Thoma, Barna Pásztor, Andreas Krause, Giorgia Ramponi 等NeurIPS 2024 · 被引用 21 次
- Aligning Individual and Collective Objectives in Multi-Agent CooperationYang Li, Wenhao Zhang, Jianhong Wang, Shao Zhang 等NeurIPS 2024 · 被引用 15 次
相关 Paper
- Reciprocal Reward Influence Encourages Cooperation From Self-Interested AgentsJohn L. Zhou, Weizhe Hong, Jonathan C. KaoNeurIPS 2024 · 被引用 5 次
- Learning Altruistic Behaviours in Reinforcement Learning without External RewardsTim Franzmeyer, Mateusz Malinowski, João F. HenriquesICLR 2022 · 被引用 10 次
- What Can Learned Intrinsic Rewards Capture?Zeyu Zheng, Junhyuk Oh, Matteo Hessel, Zhongwen Xu 等ICML 2020 · 被引用 87 次
- Multi-Agent Learning from LearnersMine Melodi Caliskan, Francesco Chini, Setareh MaghsudiICML 2023
- Learning Task-Distribution Reward Shaping with Meta-LearningHaosheng Zou, Tongzheng Ren, Dong Yan, Hang Su 等AAAI 2021 · 被引用 19 次
