Learning by Competition of Self-Interested Reinforcement Learning Agents
Stephen Chung
摘要
An artificial neural network can be trained by uniformly broadcasting a reward signal to units that implement a RE-INFORCE learning rule. Though this presents a biologically plausible alternative to backpropagation in training a network, the high variance associated with it renders it impractical to train deep networks. The high variance arises from the inefficient structural credit assignment since a single reward signal is used to evaluate the collective action of all units. To facilitate structural credit assignment, we propose replacing the reward signal to hidden units with the change in the L 2 norm of the unit's outgoing weight. As such, each hidden unit in the network is trying to maximize the norm of its outgoing weight instead of the global reward, and thus we call this learning method Weight Maximization. We prove that Weight Maximization is approximately following the gradient of rewards in expectation. In contrast to backpropagation, Weight Maximization can be used to train both continuousvalued and discrete-valued units. Moreover, Weight Maximization solves several major issues of backpropagation relating to biological plausibility. Our experiments show that a network trained with Weight Maximization can learn significantly faster than REINFORCE and slightly slower than backpropagation. Weight Maximization illustrates an example of cooperative behavior automatically arising from a population of self-interested agents in a competitive game without any central coordination.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper3
- Attention-Gated Brain Propagation: How the brain can implement reward-based error backpropagationIsabella Pozzi, Sander M. Bohté, Pieter R. RoelfsemaNeurIPS 2020 · 被引用 38 次
- Asynchronous Coagent NetworksJames E. Kostas, Chris Nota, Philip S. ThomasICML 2020 · 被引用 9 次
- MAP Propagation Algorithm: Faster Learning with a Team of Reinforcement Learning AgentsStephen ChungNeurIPS 2021 · 被引用 5 次
相关 Paper
- Learning to solve the credit assignment problemBenjamin James Lansdell, Prashanth Ravi Prakash, Konrad Paul KördingICLR 2020 · 被引用 60 次
- Structural Credit Assignment in Neural Networks using Reinforcement LearningDhawal Gupta, Gabor Mihucz, Matthew Schlegel, James E. Kostas 等NeurIPS 2021 · 被引用 9 次
- Credit Assignment Through Broadcasting a Global Error VectorDavid G. Clark, L. F. Abbott, SueYeon ChungNeurIPS 2021 · 被引用 29 次
- Learning to Learn with Feedback and Local PlasticityJack Lindsey, Ashok Litwin-KumarNeurIPS 2020 · 被引用 38 次
- Correlative Information Maximization: A Biologically Plausible Approach to Supervised Deep Neural Networks without Weight SymmetryBariscan Bozkurt, Cengiz Pehlevan, Alper T. ErdoganNeurIPS 2023 · 被引用 6 次
