Deep Innovation Protection: Confronting the Credit Assignment Problem in Training Heterogeneous Neural Architectures
Sebastian Risi, Kenneth O. Stanley
Abstract
Deep reinforcement learning approaches have shown impressive results in a variety of different domains, however, more complex heterogeneous architectures such as world models require the different neural components to be trained separately instead of end-to-end. While a simple genetic algorithm recently showed end-to-end training is possible, it failed to solve a more complex 3D task. This paper presents a method called Deep Innovation Protection (DIP) that addresses the credit assignment problem in training complex heterogenous neural network models end-to-end for such environments. The main idea behind the approach is to employ multiobjective optimization to temporally reduce the selection pressure on specific components in multi-component network, allowing other components to adapt. We investigate the emergent representations of these evolved networks, which learn to predict properties important for the survival of the agent, without the need for a specific forward-prediction loss.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 91e25ee3-a792-4978-bc5e-e5f38a33c2ecRelated papers
- A Consciousness-Inspired Planning Agent for Model-Based Reinforcement LearningMingde Zhao, Zhen Liu, Sitao Luan, Shuyuan Zhang et al.NeurIPS 2021 · 41 citations
- Deep Coordination GraphsWendelin Boehmer, Vitaly Kurin, Shimon WhitesonICML 2020 · 209 citations
- Learning Synthetic Environments and Reward Networks for Reinforcement LearningFabio Ferreira, Thomas Nierhoff, Andreas Sälinger, Frank HutterICLR 2022 · 6 citations
- Proximal Distilled Evolutionary Reinforcement LearningCristian Bodnar, Ben Day, Pietro LióAAAI 2020 · 101 citations
- OPtions as REsponses: Grounding behavioural hierarchies in multi-agent reinforcement learningAlexander Vezhnevets, Yuhuai Wu, Maria K. Eckstein, Rémi Leblond et al.ICML 2020 · 44 citations
