The Virtues of Laziness in Model-based RL: A Unified Objective and Algorithms
Anirudh Vemula, Yuda Song, Aarti Singh, Drew Bagnell, Sanjiban Choudhury
Abstract
We propose a novel approach to addressing two fundamental challenges in Model-based Reinforcement Learning (MBRL): the computational expense of repeatedly finding a good policy in the learned model, and the objective mismatch between model fitting and policy computation. Our "lazy" method leverages a novel unified objective, Performance Difference via Advantage in Model, to capture the performance difference between the learned policy and expert policy under the true dynamics. This objective demonstrates that optimizing the expected policy advantage in the learned model under an exploration distribution is sufficient for policy computation, resulting in a significant boost in computational efficiency compared to traditional planning methods. Additionally, the unified objective uses a value moment matching term for model fitting, which is aligned with the model's usage during policy computation. We present two no-regret algorithms to optimize the proposed objective, and demonstrate their statistical and computational gains compared to existing MBRL methods through simulated benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Hybrid Inverse Reinforcement LearningJuntao Ren, Gokul Swamy, Steven Wu, Drew Bagnell et al.ICML 2024 · 33 citations
- A Smooth Sea Never Made a Skilled SAILOR: Robust Imitation via Learning to SearchArnav Kumar Jain, Vibhakar Mohta, Subin Kim, Atiksh Bhardwaj et al.NeurIPS 2025 · 27 citations
- Hybrid Reinforcement Learning from Offline Observation AloneYuda Song, Drew Bagnell, Aarti SinghICML 2024 · 6 citations
- Policy-Driven World Model Adaptation for Robust Offline Model-based Reinforcement LearningJiayu Chen, Le Xu, Aravind Venugopal, Jeff SchneiderICML 2026
Builds on8
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
- Model-Based Reinforcement Learning with Value-Targeted RegressionAlex Ayoub, Zeyu Jia, Csaba Szepesvári, Mengdi Wang et al.ICML 2020 · 324 citations
- Information Theoretic Regret Bounds for Online Nonlinear ControlSham M. Kakade, Akshay Krishnamurthy, Kendall Lowrey, Motoya Ohnishi et al.NeurIPS 2020 · 137 citations
- The Value Equivalence Principle for Model-Based Reinforcement LearningChristopher Grimm, André Barreto, Satinder Singh, David SilverNeurIPS 2020 · 129 citations
- Of Moments and Matching: A Game-Theoretic Framework for Closing the Imitation GapGokul Swamy, Sanjiban Choudhury, J. Andrew Bagnell, Steven WuICML 2021 · 90 citations
Related papers
- A Unified Framework for Alternating Offline Model Training and Policy LearningShentao Yang, Shujian Zhang, Yihao Feng, Mingyuan ZhouNeurIPS 2022 · 18 citations
- How to Fine-tune the Model: Unified Model Shift and Model Bias Policy OptimizationHai Zhang, Hang Yu, Junqiao Zhao, Di Zhang et al.NeurIPS 2023 · 16 citations
- Model-based Policy Optimization with Unsupervised Model AdaptationJian Shen, Han Zhao, Weinan Zhang, Yong YuNeurIPS 2020 · 33 citations
- Ready Policy One: World Building Through Active LearningPhilip J. Ball, Jack Parker-Holder, Aldo Pacchiano, Krzysztof Choromanski et al.ICML 2020 · 52 citations
- Mismatched No More: Joint Model-Policy Optimization for Model-Based RLBenjamin Eysenbach, Alexander Khazatsky, Sergey Levine, Ruslan SalakhutdinovNeurIPS 2022 · 57 citations
