Ready Policy One: World Building Through Active Learning
Philip J. Ball, Jack Parker-Holder, Aldo Pacchiano, Krzysztof Choromanski, Stephen J. Roberts
Abstract
Model-Based Reinforcement Learning (MBRL) offers a promising direction for sample efficient learning, often achieving state of the art results for continuous control tasks. However many existing MBRL methods rely on combining greedy policies with exploration heuristics, and even those which utilize principled exploration bonuses construct dual objectives in an ad hoc fashion. In this paper we introduce Ready Policy One (RP1), a framework that views MBRL as an active learning problem, where we aim to improve the world model in the fewest samples possible. RP1 achieves this by utilizing a hybrid objective function, which crucially adapts during optimization, allowing the algorithm to trade off reward v.s. exploration at different stages of learning. In addition, we introduce a principled mechanism to terminate sample collection once we have a rich enough trajectory batch to improve the model. We rigorously evaluate our method on a variety of continuous control tasks, and demonstrate statistically significant gains over existing approaches. * Equal contribution. Preprint. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e8344b40-2d82-404b-b2b6-fb31c3e2c79bCited by top-tier papers18
- Planning to Explore via Self-Supervised World ModelsRamanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel et al.ICML 2020 · 489 citations
- Effective Diversity in Population Based Reinforcement LearningJack Parker-Holder, Aldo Pacchiano, Krzysztof Marcin Choromanski, Stephen J. RobertsNeurIPS 2020 · 195 citations
- Discovering and Achieving Goals via World ModelsRussell Mendonca, Oleh Rybkin, Kostas Daniilidis, Danijar Hafner et al.NeurIPS 2021 · 177 citations
- Tactical Optimism and Pessimism for Deep Reinforcement LearningTed Moskovitz, Jack Parker-Holder, Aldo Pacchiano, Michael Arbel et al.NeurIPS 2021 · 75 citations
- Revisiting Design Choices in Offline Model Based Reinforcement LearningCong Lu, Philip J. Ball, Jack Parker-Holder, Michael A. Osborne et al.ICLR 2022 · 65 citations
Builds on2
Related papers
- The Virtues of Laziness in Model-based RL: A Unified Objective and AlgorithmsAnirudh Vemula, Yuda Song, Aarti Singh, Drew Bagnell et al.ICML 2023 · 15 citations
- Plan To Predict: Learning an Uncertainty-Foreseeing Model For Model-Based Reinforcement LearningZifan Wu, Chao Yu, Chen Chen, Jianye Hao et al.NeurIPS 2022 · 28 citations
- A Unified Framework for Alternating Offline Model Training and Policy LearningShentao Yang, Shujian Zhang, Yihao Feng, Mingyuan ZhouNeurIPS 2022 · 18 citations
- HarmonyDream: Task Harmonization Inside World ModelsHaoyu Ma, Jialong Wu, Ningya Feng, Chenjun Xiao et al.ICML 2024 · 21 citations
- Live in the Moment: Learning Dynamics Model Adapted to Evolving PolicyXiyao Wang, Wichayaporn Wongkamjan, Ruonan Jia, Furong HuangICML 2023 · 20 citations
