Any-step Dynamics Model Improves Future Predictions for Online and Offline Reinforcement Learning
Haoxin Lin, Yu-Yan Xu, Yihao Sun, Zhilong Zhang, Yi-Chen Li, Chengxing Jia, Junyin Ye, Jiaji Zhang, Yang Yu
Abstract
Model-based methods in reinforcement learning offer a promising approach to enhance data efficiency by facilitating policy exploration within a dynamics model. However, accurately predicting sequential steps in the dynamics model remains a challenge due to the bootstrapping prediction, which attributes the next state to the prediction of the current state. This leads to accumulated errors during model roll-out. In this paper, we propose the Any-step Dynamics Model (ADM) to mitigate the compounding error by reducing bootstrapping prediction to direct prediction. ADM allows for the use of variable-length plans as inputs for predicting future states without frequent bootstrapping. We design two algorithms, ADMPO-ON and ADMPO-OFF, which apply ADM in online and offline modelbased frameworks, respectively. In the online setting, ADMPO-ON demonstrates improved sample efficiency compared to previous state-of-the-art methods. In the offline setting, ADMPO-OFF not only demonstrates superior performance compared to recent state-of-the-art offline approaches but also offers better quantification of model uncertainty using only a single ADM. The code is available at https://github.com/LAMDA-RL/ADMPO.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ba0c8dbd-c907-40b6-8c01-cae8e5b55700Cited by top-tier papers11
- Model-based Offline RL via Robust Value-Aware Model Learning with Implicitly Differentiable Adaptive WeightingZhongjian Qiao, Jiafei Lyu, Boxiang Lyu, Yao Shu et al.ICLR 2026 · 5 citations
- Towards Practical World Model-based Reinforcement Learning for Vision-Language-Action ModelsZhilong Zhang, Haoxiang Ren, Yihao Sun, Yifei Sheng et al.ICML 2026 · 3 citations
- Regularized Offline Policy Optimization with Posterior Hybrid Bayesian BeliefHongqiang Lin, Pengfei Wang, Nenggan ZhengICML 2026 · 1 citation
- Offline Reinforcement Learning with Universal Horizon ModelsHojun Chung, Junseo Lee, Songhwai OhICML 2026 · 1 citation
- Long-Horizon Model-Based Offline Reinforcement Learning Without Explicit ConservatismTianwei Ni, Esther Derman, Vineet Jain, Vincent Taboga et al.ICML 2026 · 1 citation
Builds on24
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 870 citations
- COMBO: Conservative Offline Model-Based Policy OptimizationTianhe Yu, Aviral Kumar, Rafael Rafailov, Aravind Rajeswaran et al.NeurIPS 2021 · 549 citations
Related papers
- ADM-v2: Pursuing Full-Horizon Roll-out in Dynamics Models for Offline Policy Learning and EvaluationHaoxin Lin, Siyuan Xiao, Yi-Chen Li, Zhilong Zhang et al.ICLR 2026
- Bidirectional Model-based Policy OptimizationHang Lai, Jian Shen, Weinan Zhang, Yong YuICML 2020 · 66 citations
- On Rollouts in Model-Based Reinforcement LearningBernd Frauenknecht, Devdutt Subhasish, Friedrich Solowjow, Sebastian TrimpeICLR 2025 · 1 citation
- COPlanner: Plan to Roll Out Conservatively but to Explore Optimistically for Model-Based RLXiyao Wang, Ruijie Zheng, Yanchao Sun, Ruonan Jia et al.ICLR 2024 · 19 citations
- Diminishing Return of Value Expansion Methods in Model-Based Reinforcement LearningDaniel Palenicek, Michael Lutter, Joao Carvalho, Jan PetersICLR 2023
