Frequency-based Search-control in Dyna
Yangchen Pan, Jincheng Mei, Amir-massoud Farahmand
摘要
Model-based reinforcement learning has been empirically demonstrated as a successful strategy to improve sample efficiency. In particular, Dyna is an elegant model-based architecture integrating learning and planning that provides huge flexibility of using a model. One of the most important components in Dyna is called search-control, which refers to the process of generating state or state-action pairs from which we query the model to acquire simulated experiences. Search-control is critical in improving learning efficiency. In this work, we propose a simple and novel search-control strategy by searching high frequency regions of the value function. Our main intuition is built on Shannon sampling theorem from signal processing, which indicates that a high frequency signal requires more samples to reconstruct. We empirically show that a high frequency function is more difficult to approximate. This suggests a search-control strategy: we should use states from high frequency regions of the value function to query the model to acquire more samples. We develop a simple strategy to locally measure the frequency of a function by gradient and hessian norms, and provide theoretical justification for this approach. We then apply our strategy to search-control in Dyna, and conduct experiments to show its property and effectiveness on benchmark domains.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Model-augmented Prioritized Experience ReplayYoungmin Oh, Jinwoo Shin, Eunho Yang, Sung Ju HwangICLR 2022 · 被引用 21 次
- Learning to Sample with Local and Global Contexts in Experience Replay BufferYoungmin Oh, Kimin Lee, Jinwoo Shin, Eunho Yang 等ICLR 2021 · 被引用 19 次
- Curious Replay for Model-based AdaptationIsaac Kauvar, Chris Doyle, Linqi Zhou, Nick HaberICML 2023 · 被引用 18 次
- Self-Consistent Models and ValuesGregory Farquhar, Kate Baumli, Zita Marinho, Angelos Filos 等NeurIPS 2021 · 被引用 10 次
- MAD-TD: Model-Augmented Data stabilizes High Update Ratio RLClaas Voelcker, Marcel Hussing, Eric Eaton, Amir-massoud Farahmand 等ICLR 2025
它引用的顶会 Paper1
相关 Paper
- Live in the Moment: Learning Dynamics Model Adapted to Evolving PolicyXiyao Wang, Wichayaporn Wongkamjan, Ruonan Jia, Furong HuangICML 2023 · 被引用 20 次
- COPlanner: Plan to Roll Out Conservatively but to Explore Optimistically for Model-Based RLXiyao Wang, Ruijie Zheng, Yanchao Sun, Ruonan Jia 等ICLR 2024 · 被引用 19 次
- An Experimental Design Perspective on Model-Based Reinforcement LearningViraj Mehta, Biswajit Paria, Jeff Schneider, Stefano Ermon 等ICLR 2022 · 被引用 24 次
- On Effective Scheduling of Model-based Reinforcement LearningHang Lai, Jian Shen, Weinan Zhang, Yimin Huang 等NeurIPS 2021 · 被引用 23 次
- Sample-Efficient Reinforcement Learning via Conservative Model-Based Actor-CriticZhihai Wang, Jie Wang, Qi Zhou, Bin Li 等AAAI 2022 · 被引用 38 次
