Dynamic Update-to-Data Ratio: Minimizing World Model Overfitting
Nicolai Dorka, Tim Welschehold, Wolfram Burgard
摘要
Early stopping based on the validation set performance is a popular approach to find the right balance between under- and overfitting in the context of supervised learning. However, in reinforcement learning, even for supervised sub-problems such as world model learning, early stopping is not applicable as the dataset is continually evolving. As a solution, we propose a new general method that dynamically adjusts the update to data (UTD) ratio during training based on under- and overfitting detection on a small subset of the continuously collected experience not used for training. We apply our method to DreamerV2, a state-of-the-art model-based reinforcement learning algorithm, and evaluate it on the DeepMind Control Suite and the Atari k benchmark. The results demonstrate that one can better balance under- and overestimation by adjusting the UTD ratio with our approach compared to the default setting in DreamerV2 and that it is competitive with an extensive hyperparameter search which is not feasible for many applications. Our method eliminates the need to set the UTD hyperparameter by hand and even leads to a higher robustness with regard to other learning-related hyperparameters further reducing the amount of necessary tuning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- SUF: Stabilized Unconstrained Fine-Tuning for Offline-to-Online Reinforcement LearningJiaheng Feng, Mingxiao Feng, Haolin Song, Wengang Zhou 等AAAI 2024 · 被引用 7 次
- Langevin Soft Actor-Critic: Efficient Exploration through Uncertainty-Driven Critic LearningHaque Ishfaq, Guangyuan Wang, Sami Nur Islam, Doina PrecupICLR 2025
它引用的顶会 Paper9
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 被引用 1,170 次
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
- Model Based Reinforcement Learning for AtariLukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski 等ICLR 2020 · 被引用 969 次
- Reinforcement Learning with Augmented DataMichael Laskin, Kimin Lee, Adam Stooke, Lerrel Pinto 等NeurIPS 2020 · 被引用 833 次
相关 Paper
- Compute-Optimal Scaling for Value-Based Deep RLPreston Fu, Oleh Rybkin, Zhiyuan Zhou, Michal Nauman 等NeurIPS 2025 · 被引用 7 次
- DyMoDreamer: World Modeling with Dynamic ModulationBoxuan Zhang, Runqing Wang, Wei Xiao, Weipu Zhang 等NeurIPS 2025 · 被引用 2 次
- MAD-TD: Model-Augmented Data stabilizes High Update Ratio RLClaas Voelcker, Marcel Hussing, Eric Eaton, Amir-massoud Farahmand 等ICLR 2025
- Value-Based Deep RL Scales PredictablyOleh Rybkin, Michal Nauman, Preston Fu, Charlie Victor Snell 等ICML 2025
- Curious Replay for Model-based AdaptationIsaac Kauvar, Chris Doyle, Linqi Zhou, Nick HaberICML 2023 · 被引用 18 次
