Learning Synthetic Environments and Reward Networks for Reinforcement Learning
Fabio Ferreira, Thomas Nierhoff, Andreas Sälinger, Frank Hutter
摘要
We introduce Synthetic Environments (SEs) and Reward Networks (RNs), represented by neural networks, as proxy environment models for training Reinforcement Learning (RL) agents. We show that an agent, after being trained exclusively on the SE, is able to solve the corresponding real environment. While an SE acts as a full proxy to a real environment by learning about its state dynamics and rewards, an RN is a partial proxy that learns to augment or replace rewards. We use bi-level optimization to evolve SEs and RNs: the inner loop trains the RL agent, and the outer loop trains the parameters of the SE / RN via an evolution strategy. We evaluate our proposed new concept on a broad range of RL algorithms and classic control environments. In a one-to-one comparison, learning an SE proxy requires more interactions with the real environment than training agents only on the real environment. However, once such an SE has been learned, we do not need any interactions with the real environment to train new agents. Moreover, the learned SE proxies allow us to train agents with fewer interactions while maintaining the original task performance. Our empirical results suggest that SEs achieve this result by learning informed representations that bias the agents towards relevant states. Moreover, we find that these proxies are robust against hyperparameter variation and can also transfer to unseen agents.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper5
- Leveraging Procedural Generation to Benchmark Reinforcement LearningKarl Cobbe, Christopher Hesse, Jacob Hilton, John SchulmanICML 2020 · 被引用 685 次
- Emergent Complexity and Zero-shot Transfer via Unsupervised Environment DesignMichael Dennis, Natasha Jaques, Eugene Vinitsky, Alexandre M. Bayen 等NeurIPS 2020 · 被引用 362 次
- Learning to Utilize Shaping Rewards: A New Approach of Reward ShapingYujing Hu, Weixun Wang, Hangtian Jia, Yixiang Wang 等NeurIPS 2020 · 被引用 256 次
- Generative Teaching Networks: Accelerating Neural Architecture Search by Learning to Generate Synthetic Training DataFelipe Petroski Such, Aditya Rawal, Joel Lehman, Kenneth O. Stanley 等ICML 2020 · 被引用 180 次
- What Can Learned Intrinsic Rewards Capture?Zeyu Zheng, Junhyuk Oh, Matteo Hessel, Zhongwen Xu 等ICML 2020 · 被引用 87 次
相关 Paper
- Evolutionary Reinforcement Learning for Sample-Efficient Multiagent CoordinationSomdeb Majumdar, Shauharda Khadka, Santiago Miret, Stephen McAleer 等ICML 2020 · 被引用 70 次
- Deep Innovation Protection: Confronting the Credit Assignment Problem in Training Heterogeneous Neural ArchitecturesSebastian Risi, Kenneth O. StanleyAAAI 2021
- Learning Human Objectives by Evaluating Hypothetical BehaviorSiddharth Reddy, Anca D. Dragan, Sergey Levine, Shane Legg 等ICML 2020 · 被引用 81 次
- Safe and Scalable Web Agent Learning via Recreated WebsitesHyungjoo Chae, Jungsoo Park, Alan RitterICML 2026
- Improving Generalization in Meta Reinforcement Learning using Learned ObjectivesLouis Kirsch, Sjoerd van Steenkiste, Jürgen SchmidhuberICLR 2020 · 被引用 132 次
