Is Model Ensemble Necessary? Model-based RL via a Single Model with Lipschitz Regularized Value Function
Ruijie Zheng, Xiyao Wang, Huazhe Xu, Furong Huang
Abstract
Probabilistic dynamics model ensemble is widely used in existing model-based reinforcement learning methods as it outperforms a single dynamics model in both asymptotic performance and sample efficiency. In this paper, we provide both practical and theoretical insights on the empirical success of the probabilistic dynamics model ensemble through the lens of Lipschitz continuity. We find that, for a value function, the stronger the Lipschitz condition is, the smaller the gap between the true dynamics-and learned dynamics-induced Bellman operators is, thus enabling the converged value function to be closer to the optimal value function. Hence, we hypothesize that the key functionality of the probabilistic dynamics model ensemble is to regularize the Lipschitz condition of the value function using generated samples. To test this hypothesis, we devise two practical robust training mechanisms through computing the adversarial noise and regularizing the value network's spectral norm to directly regularize the Lipschitz condition of the value functions. Empirical results show that combined with our mechanisms, model-based RL algorithms with a single dynamics model outperform those with an ensemble of probabilistic dynamics models. These findings not only support the theoretical insight, but also provide a practical solution for developing computationally efficient model-based RL algorithms. § Equal contribution This value-aware model error measures the difference between the simulated and true Bellman operator acting on a given Q function. In other words, even though the learned transition model Pθ
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1632ea38-6ef4-469d-bc64-4ed079b734f5Cited by top-tier papers9
- Value Gradient weighted Model-Based Reinforcement LearningClaas Voelcker, Victor Liao, Animesh Garg, Amir-massoud FarahmandICLR 2022 · 37 citations
- Live in the Moment: Learning Dynamics Model Adapted to Evolving PolicyXiyao Wang, Wichayaporn Wongkamjan, Ruonan Jia, Furong HuangICML 2023 · 20 citations
- COPlanner: Plan to Roll Out Conservatively but to Explore Optimistically for Model-Based RLXiyao Wang, Ruijie Zheng, Yanchao Sun, Ruonan Jia et al.ICLR 2024 · 19 citations
- PRISE: LLM-Style Sequence Compression for Learning Temporal Action Abstractions in ControlRuijie Zheng, Ching-An Cheng, Hal Daumé III, Furong Huang et al.ICML 2024 · 17 citations
- Physics-informed Value Learner for Offline Goal-Conditioned Reinforcement LearningVittorio Giammarino, Ruiqi Ni, Ahmed H. QureshiNeurIPS 2025 · 15 citations
Builds on11
- Robust Deep Reinforcement Learning against Adversarial Perturbations on State ObservationsHuan Zhang, Hongge Chen, Chaowei Xiao, Bo Li et al.NeurIPS 2020 · 437 citations
- The Value Equivalence Principle for Model-Based Reinforcement LearningChristopher Grimm, André Barreto, Satinder Singh, David SilverNeurIPS 2020 · 129 citations
- Model-Augmented Actor-Critic: Backpropagating through PathsIgnasi Clavera, Yao Fu, Pieter AbbeelICLR 2020 · 96 citations
- Deep Reinforcement Learning with Robust and Smooth PolicyQianli Shen, Yan Li, Haoming Jiang, Zhaoran Wang et al.ICML 2020 · 95 citations
- Spectral Normalisation for Deep Reinforcement Learning: An Optimisation PerspectiveFlorin Gogianu, Tudor Berariu, Mihaela Rosca, Claudia Clopath et al.ICML 2021 · 69 citations
Related papers
- Model-based Reinforcement Learning for Parameterized Action SpacesRenhao Zhang, Haotian Fu, Yilin Miao, George KonidarisICML 2024 · 8 citations
- Adversarial Diffusion for Robust Reinforcement LearningDaniele Foffano, Alessio Russo, Alexandre ProutièreNeurIPS 2025 · 5 citations
- Prioritized Model Experience ReplayMuxi Tao, jiangtao wen, Yuxing HanICML 2026
- Robust Reinforcement Learning for Continuous Control with Model MisspecificationDaniel J. Mankowitz, Nir Levine, Rae Jeong, Abbas Abdolmaleki et al.ICLR 2020 · 138 citations
- Robust Model Based Reinforcement Learning Using L1 Adaptive ControlMinjun Sung, Sambhu H. Karumanchi, Aditya Gahlawat, Naira HovakimyanICLR 2024 · 1 citation
