MAP Propagation Algorithm: Faster Learning with a Team of Reinforcement Learning Agents
Stephen Chung
Abstract
Nearly all state-of-the-art deep learning algorithms rely on error backpropagation, which is generally regarded as biologically implausible. An alternative way of training an artificial neural network is through treating each unit in the network as a reinforcement learning agent, and thus the network is considered as a team of agents. As such, all units can be trained by REINFORCE, a local learning rule modulated by a global signal that is more consistent with biologically observed forms of synaptic plasticity. Although this learning rule follows the gradient of return in expectation, it suffers from high variance and thus the low speed of learning, rendering it impractical to train deep networks. We therefore propose a novel algorithm called MAP propagation to reduce this variance significantly while retaining the local property of the learning rule. Experiments demonstrated that MAP propagation could solve common reinforcement learning tasks at a similar speed to backpropagation when applied to an actor-critic network. Our work thus allows for the broader application of the teams of agents in deep reinforcement learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5ef28586-186b-4d52-bb4f-d172da38141fCited by top-tier papers1
Ask how each one uses itBuilds on1
Related papers
- Learning to solve the credit assignment problemBenjamin James Lansdell, Prashanth Ravi Prakash, Konrad Paul KördingICLR 2020 · 60 citations
- Asynchronous Coagent NetworksJames E. Kostas, Chris Nota, Philip S. ThomasICML 2020 · 9 citations
- Can the Brain Do Backpropagation? - Exact Implementation of Backpropagation in Predictive Coding NetworksYuhang Song, Thomas Lukasiewicz, Zhenghua Xu, Rafal BogaczNeurIPS 2020 · 117 citations
- Dendritic Localized Learning: Toward Biologically Plausible AlgorithmChangze Lv, Jingwen Xu, Yiyang Lu, Xiaohua Wang et al.ICML 2025
- GAIT-prop: A biologically plausible learning rule derived from backpropagation of errorNasir Ahmad, Marcel A. J. van Gerven, Luca AmbrogioniNeurIPS 2020 · 28 citations
