Scalable Offline Model-Based RL with Action Chunks
Kwanyoung Park, Seohong Park, Youngwoon Lee, Sergey Levine
摘要
In this paper, we study whether model-based reinforcement learning (RL), in particular model-based value expansion, can provide a scalable recipe for tackling complex, long-horizon tasks in offline RL. Model-based value expansion fits an on-policy value function using length-n imaginary rollouts generated by the current policy and a learned dynamics model. While larger n reduces bias in value bootstrapping, it amplifies accumulated model errors over long horizons, degrading future predictions. We address this trade-off with an action-chunk model that predicts a future state from a sequence of actions (an"action chunk") instead of a single action, which reduces compounding errors. In addition, instead of directly training a policy to maximize rewards, we employ rejection sampling from an expressive behavioral action-chunk policy, which prevents model exploitation from out-of-distribution actions. We call this recipe Model-Based RL with Action Chunks (MAC). Through experiments on highly challenging tasks with large-scale datasets of up to 100M transitions, we show that MAC achieves the best performance among offline model-based RL algorithms, especially on challenging long-horizon tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Towards Practical World Model-based Reinforcement Learning for Vision-Language-Action ModelsZhilong Zhang, Haoxiang Ren, Yihao Sun, Yifei Sheng 等ICML 2026 · 被引用 3 次
- Chunk-Guided Q-LearningGwanwoo Song, Kwanyoung Park, Youngwoon LeeICML 2026 · 被引用 2 次
- Offline Reinforcement Learning with Universal Horizon ModelsHojun Chung, Junseo Lee, Songhwai OhICML 2026 · 被引用 1 次
- Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-LearningSungyoung Lee, Dohyeong Kim, Eshan Balachandar, Zelal Mustafaoglu 等ICML 2026
它引用的顶会 Paper43
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee 等NeurIPS 2021 · 被引用 2,557 次
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
相关 Paper
- Reinforcement Learning with Action ChunkingQiyang Li, Zhiyuan Zhou, Sergey LevineNeurIPS 2025 · 被引用 114 次
- Decoupled Q-ChunkingQiyang Li, Seohong Park, Sergey LevineICLR 2026 · 被引用 19 次
- The Edge-of-Reach Problem in Offline Model-Based Reinforcement LearningAnya Sims, Cong Lu, Jakob N. Foerster, Yee Whye TehNeurIPS 2024 · 被引用 9 次
- Model-based Offline Reinforcement Learning with Lower Expectile Q-LearningKwanyoung Park, Youngwoon LeeICLR 2025
- Diminishing Return of Value Expansion Methods in Model-Based Reinforcement LearningDaniel Palenicek, Michael Lutter, Joao Carvalho, Jan PetersICLR 2023
