Mega-Reward: Achieving Human-Level Play without Extrinsic Rewards
Yuhang Song, Jianyi Wang, Thomas Lukasiewicz, Zhenghua Xu, Shangtong Zhang, Andrzej Wojcicki, Mai Xu
摘要
Intrinsic rewards were introduced to simulate how human intelligence works; they are usually evaluated by intrinsically-motivated play, i.e., playing games without extrinsic rewards but evaluated with extrinsic rewards. However, none of the existing intrinsic reward approaches can achieve human-level performance under this very challenging setting of intrinsically-motivated play. In this work, we propose a novel megalomania-driven intrinsic reward (called mega-reward), which, to our knowledge, is the first approach that achieves human-level performance in intrinsically-motivated play. Intuitively, mega-reward comes from the observation that infants' intelligence develops when they try to gain more control on entities in an environment; therefore, mega-reward aims to maximize the control capabilities of agents on given entities in a given environment. To formalize mega-reward, a relational transition model is proposed to bridge the gaps between direct and latent control. Experimental studies show that mega-reward (i) can greatly outperform all state-of-the-art intrinsic reward approaches, (ii) generally achieves the same level of performance as Ex-PPO and professional human-level scores, and (iii) has also a superior performance when it is incorporated with extrinsic rewards.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Causal Influence Detection for Improving Efficiency in Reinforcement LearningMaximilian Seitzer, Bernhard Schölkopf, Georg MartiusNeurIPS 2021 · 被引用 120 次
- Mutual Information State Intrinsic ControlRui Zhao, Yang Gao, Pieter Abbeel, Volker Tresp 等ICLR 2021 · 被引用 25 次
- Causal Action Influence Aware Counterfactual Data AugmentationNúria Armengol Urpí, Marco Bagatella, Marin Vlastelica, Georg MartiusICML 2024 · 被引用 11 次
相关 Paper
- Redeeming intrinsic rewards via constrained optimizationEric Chen, Zhang-Wei Hong, Joni Pajarinen, Pulkit AgrawalNeurIPS 2022 · 被引用 48 次
- Regularity as Intrinsic Reward for Free PlayCansu Sancaktar, Justus H. Piater, Georg MartiusNeurIPS 2023 · 被引用 9 次
- Motif: Intrinsic Motivation from Artificial Intelligence FeedbackMartin Klissarov, Pierluca D'Oro, Shagun Sodhani, Roberta Raileanu 等ICLR 2024 · 被引用 97 次
- Learning with AMIGo: Adversarially Motivated Intrinsic GoalsAndres Campero, Roberta Raileanu, Heinrich Küttler, Joshua B. Tenenbaum 等ICLR 2021 · 被引用 48 次
- Guiding Pretraining in Reinforcement Learning with Large Language ModelsYuqing Du, Olivia Watkins, Zihan Wang, Cédric Colas 等ICML 2023 · 被引用 257 次
