Learning by Competition of Self-Interested Reinforcement Learning Agents
Stephen Chung
Abstract
An artificial neural network can be trained by uniformly broadcasting a reward signal to units that implement a RE-INFORCE learning rule. Though this presents a biologically plausible alternative to backpropagation in training a network, the high variance associated with it renders it impractical to train deep networks. The high variance arises from the inefficient structural credit assignment since a single reward signal is used to evaluate the collective action of all units. To facilitate structural credit assignment, we propose replacing the reward signal to hidden units with the change in the L 2 norm of the unit's outgoing weight. As such, each hidden unit in the network is trying to maximize the norm of its outgoing weight instead of the global reward, and thus we call this learning method Weight Maximization. We prove that Weight Maximization is approximately following the gradient of rewards in expectation. In contrast to backpropagation, Weight Maximization can be used to train both continuousvalued and discrete-valued units. Moreover, Weight Maximization solves several major issues of backpropagation relating to biological plausibility. Our experiments show that a network trained with Weight Maximization can learn significantly faster than REINFORCE and slightly slower than backpropagation. Weight Maximization illustrates an example of cooperative behavior automatically arising from a population of self-interested agents in a competitive game without any central coordination.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4ef35559-406b-48a3-ac76-e0a93559aec5Cited by top-tier papers1
Ask how each one uses itBuilds on3
- Attention-Gated Brain Propagation: How the brain can implement reward-based error backpropagationIsabella Pozzi, Sander M. Bohté, Pieter R. RoelfsemaNeurIPS 2020 · 38 citations
- Asynchronous Coagent NetworksJames E. Kostas, Chris Nota, Philip S. ThomasICML 2020 · 9 citations
- MAP Propagation Algorithm: Faster Learning with a Team of Reinforcement Learning AgentsStephen ChungNeurIPS 2021 · 5 citations
Related papers
- Learning to solve the credit assignment problemBenjamin James Lansdell, Prashanth Ravi Prakash, Konrad Paul KördingICLR 2020 · 60 citations
- Structural Credit Assignment in Neural Networks using Reinforcement LearningDhawal Gupta, Gabor Mihucz, Matthew Schlegel, James E. Kostas et al.NeurIPS 2021 · 9 citations
- Credit Assignment Through Broadcasting a Global Error VectorDavid G. Clark, L. F. Abbott, SueYeon ChungNeurIPS 2021 · 29 citations
- Learning to Learn with Feedback and Local PlasticityJack Lindsey, Ashok Litwin-KumarNeurIPS 2020 · 38 citations
- Correlative Information Maximization: A Biologically Plausible Approach to Supervised Deep Neural Networks without Weight SymmetryBariscan Bozkurt, Cengiz Pehlevan, Alper T. ErdoganNeurIPS 2023 · 6 citations
