Decentralized Reinforcement Learning: Global Decision-Making via Local Economic Transactions
Michael Chang, Sidhant Kaushik, S. Matthew Weinberg, Tom Griffiths, Sergey Levine
摘要
This paper 1 seeks to establish a framework for directing a society of simple, specialized, selfinterested agents to solve what traditionally are posed as monolithic single-agent sequential decision problems. What makes it challenging to use a decentralized approach to collectively optimize a central objective is the difficulty in characterizing the equilibrium strategy profile of noncooperative games. To overcome this challenge, we design a mechanism for defining the learning environment of each agent for which we know that the optimal solution for the global objective coincides with a Nash equilibrium strategy profile of the agents optimizing their own local objectives. The society functions as an economy of agents that learn the credit assignment process itself by buying and selling to each other the right to operate on the environment state. We derive a class of decentralized reinforcement learning algorithms that are broadly applicable not only to standard reinforcement learning but also for selecting options in semi-MDPs and dynamically composing computation graphs. Lastly, we demonstrate the potential advantages of a society's inherent modular structure for more efficient transfer learning. You know that everything you think and do is thought and done by you. But what's a "you"? What kinds of smaller entities cooperate inside your mind to do your work? (Minsky, 1988
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- The Sensory Neuron as a Transformer: Permutation-Invariant Neural Networks for Reinforcement LearningYujin Tang, David HaNeurIPS 2021 · 被引用 90 次
- Automated Design of Affine Maximizer Mechanisms in Dynamic SettingsMichael J. Curry, Vinzenz Thoma, Darshan Chakrabarti, Stephen McAleer 等AAAI 2024 · 被引用 13 次
- MSRL: Distributed Reinforcement Learning with Dataflow FragmentsHuanzhou Zhu, Bo Zhao, Gang Chen, Weifeng Chen 等USENIX ATC 2023 · 被引用 9 次
- Modularity in Reinforcement Learning via Algorithmic Independence in Credit AssignmentMichael Chang, Sidhant Kaushik, Sergey Levine, Tom GriffithsICML 2021 · 被引用 8 次
- Sparse Distributed Memory is a Continual LearnerTrenton Bricken, Xander Davies, Deepak Singh, Dmitry Krotov 等ICLR 2023 · 被引用 5 次
它引用的顶会 Paper1
相关 Paper
- Specification-Guided Learning of Nash Equilibria with High Social WelfareKishor Jothimurugan, Suguman Bansal, Osbert Bastani, Rajeev AlurCAV 2022 · 被引用 9 次
- Stackelberg Learning with Outcome-based PaymentTom Yan, Chicheng ZhangNeurIPS 2025
- LOQA: Learning with Opponent Q-Learning AwarenessMilad Aghajohari, Juan Agustin Duque, Tim Cooijmans, Aaron C. CourvilleICLR 2024 · 被引用 9 次
- Inducing Equilibria via Incentives: Simultaneous Design-and-Play Ensures Global ConvergenceBoyi Liu, Jiayang Li, Zhuoran Yang, Hoi-To Wai 等NeurIPS 2022 · 被引用 29 次
- Sample-Efficient Multi-Agent RL: An Optimization PerspectiveNuoya Xiong, Zhihan Liu, Zhaoran Wang, Zhuoran YangICLR 2024 · 被引用 2 次
