Self-Consistent Models and Values
Gregory Farquhar, Kate Baumli, Zita Marinho, Angelos Filos, Matteo Hessel, Hado Philip van Hasselt, David Silver
摘要
Learned models of the environment provide reinforcement learning (RL) agents with flexible ways of making predictions about the environment. In particular, models enable planning, i.e. using more computation to improve value functions or policies, without requiring additional environment interactions. In this work, we investigate a way of augmenting model-based RL, by additionally encouraging a learned model and value function to be jointly self-consistent. Our approach differs from classic planning methods such as Dyna, which only update values to be consistent with the model. We propose multiple self-consistency updates, evaluate these in both tabular and function approximation settings, and find that, with appropriate choices, self-consistency helps both policy evaluation and control.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Human-Timescale Adaptation in an Open-Ended Task SpaceJakob Bauer, Kate Baumli, Feryal M. P. Behbahani, Avishkar Bhoopchand 等ICML 2023 · 被引用 155 次
- Value-Consistent Representation Learning for Data-Efficient Reinforcement LearningYang Yue, Bingyi Kang, Zhongwen Xu, Gao Huang 等AAAI 2023 · 被引用 19 次
- Model-Value Inconsistency as a Signal for Epistemic UncertaintyAngelos Filos, Eszter Vértes, Zita Marinho, Gregory Farquhar 等ICML 2022 · 被引用 9 次
- Approximate Value EquivalenceChristopher Grimm, André Barreto, Satinder SinghNeurIPS 2022 · 被引用 7 次
- Generalized Weighted Path Consistency for Mastering Atari GamesDengwei Zhao, Shikui Tu, Lei XuNeurIPS 2023 · 被引用 5 次
它引用的顶会 Paper10
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- Model Based Reinforcement Learning for AtariLukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski 等ICLR 2020 · 被引用 969 次
- Stochastic Latent Actor-Critic: Deep Reinforcement Learning with a Latent Variable ModelAlex X. Lee, Anusha Nagabandi, Pieter Abbeel, Sergey LevineNeurIPS 2020 · 被引用 437 次
- Online and Offline Reinforcement Learning by Planning with a Learned ModelJulian Schrittwieser, Thomas Hubert, Amol Mandhane, Mohammadamin Barekatain 等NeurIPS 2021 · 被引用 149 次
- The Value Equivalence Principle for Model-Based Reinforcement LearningChristopher Grimm, André Barreto, Satinder Singh, David SilverNeurIPS 2020 · 被引用 129 次
相关 Paper
- COPlanner: Plan to Roll Out Conservatively but to Explore Optimistically for Model-Based RLXiyao Wang, Ruijie Zheng, Yanchao Sun, Ruonan Jia 等ICLR 2024 · 被引用 19 次
- Frequency-based Search-control in DynaYangchen Pan, Jincheng Mei, Amir-massoud FarahmandICLR 2020 · 被引用 16 次
- Simplified Temporal Consistency Reinforcement LearningYi Zhao, Wenshuai Zhao, Rinu Boney, Juho Kannala 等ICML 2023 · 被引用 19 次
- Operator Splitting Value IterationAmin Rakhsha, Andrew Wang, Mohammad Ghavamzadeh, Amir-massoud FarahmandNeurIPS 2022 · 被引用 11 次
- Autoregressive Dynamics Models for Offline Policy Evaluation and OptimizationMichael R. Zhang, Thomas Paine, Ofir Nachum, Cosmin Paduraru 等ICLR 2021 · 被引用 52 次
