How Should an Agent Practice?
Janarthanan Rajendran, Richard L. Lewis, Vivek Veeriah, Honglak Lee, Satinder Singh
Abstract
We present a method for learning intrinsic reward functions to drive the learning of an agent during periods of practice in which extrinsic task rewards are not available. During practice, the environment may differ from the one available for training and evaluation with extrinsic rewards. We refer to this setup of alternating periods of practice and objective evaluation as practice-match, drawing an analogy to regimes of skill acquisition common for humans in sports and games. The agent must effectively use periods in the practice environment so that performance improves during matches. In the proposed method the intrinsic practice reward is learned through a meta-gradient approach that adapts the practice reward parameters to reduce the extrinsic match reward loss computed from matches. We illustrate the method on a simple grid world, and evaluate it in two games in which the practice environment differs from match: Pong with practice against a wall without an opponent, and PacMan with practice in a maze without ghosts. The results show gains from learning in practice in addition to match periods over learning in matches only.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b79dbaf7-d35d-4d5b-80bb-046eefed454aCited by top-tier papers3
- Interesting Object, Curious Agent: Learning Task-Agnostic ExplorationSimone Parisi, Victoria Dean, Deepak Pathak, Abhinav GuptaNeurIPS 2021 · 58 citations
- Unsupervised Reinforcement Learning in Multiple EnvironmentsMirco Mutti, Mattia Mancassola, Marcello RestelliAAAI 2022 · 30 citations
- Intelligent Switching for Reset-Free RLDarshan Patil, Janarthanan Rajendran, Glen Berseth, Sarath ChandarICLR 2024 · 1 citation
Related papers
- What Can Learned Intrinsic Rewards Capture?Zeyu Zheng, Junhyuk Oh, Matteo Hessel, Zhongwen Xu et al.ICML 2020 · 87 citations
- Learning with AMIGo: Adversarially Motivated Intrinsic GoalsAndres Campero, Roberta Raileanu, Heinrich Küttler, Joshua B. Tenenbaum et al.ICLR 2021 · 48 citations
- Discovery of Options via Meta-Learned SubgoalsVivek Veeriah, Tom Zahavy, Matteo Hessel, Zhongwen Xu et al.NeurIPS 2021 · 38 citations
- MetaCURE: Meta Reinforcement Learning with Empowerment-Driven ExplorationJin Zhang, Jianhao Wang, Hao Hu, Tong Chen et al.ICML 2021 · 33 citations
- Redeeming intrinsic rewards via constrained optimizationEric Chen, Zhang-Wei Hong, Joni Pajarinen, Pulkit AgrawalNeurIPS 2022 · 48 citations
