Scaling Offline Model-Based RL via Jointly-Optimized World-Action Model Pretraining
Jie Cheng, Ruixi Qiao, Yingwei Ma, Binhua Li, Gang Xiong, Qinghai Miao, Yongbin Li, Yisheng Lv
摘要
A significant aspiration of offline reinforcement learning (RL) is to develop a generalist agent with high capabilities from large and heterogeneous datasets. However, prior approaches that scale offline RL either rely heavily on expert trajectories or struggle to generalize to diverse unseen tasks. Inspired by the excellent generalization of world model in conditional video generation, we explore the potential of image observation-based world model for scaling offline RL and enhancing generalization on novel tasks. In this paper, we introduce JOWA: Jointly-Optimized World-Action model, an offline model-based RL agent pretrained on multiple Atari games with 6 billion tokens data to learn general-purpose representation and decision-making ability. Our method jointly optimizes a world-action model through a shared transformer backbone, which stabilize temporal difference learning with large models during pretraining. Moreover, we propose a provably efficient and parallelizable planning algorithm to compensate for the Q-value estimation error and thus search out better policies. Experimental results indicate that our largest agent, with 150 million parameters, achieves 78.9% human-level performance on pretrained games using only 10% subsampled offline data, outperforming existing state-of-the-art large-scale offline RL baselines by 31.6% on averange. Furthermore, JOWA scales favorably with model capacity and can sample-efficiently transfer to novel games using only 5k offline fine-tuning data (approximately 4 trajectories) per game, demonstrating superior generalization. We will release codes and model weights at https://github.com/CJReinforce/JOWA
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Horizon Reduction Makes RL ScalableSeohong Park, Kevin Frans, Deepinder Mann, Benjamin Eysenbach 等NeurIPS 2025 · 被引用 60 次
- Stop Summation: Min-Form Credit Assignment Is All Process Reward Model Needs for ReasoningJie Cheng, Gang Xiong, Ruixi Qiao, Lijun Li 等NeurIPS 2025 · 被引用 56 次
- Scalable Offline Model-Based RL with Action ChunksKwanyoung Park, Seohong Park, Youngwoon Lee, Sergey LevineICLR 2026 · 被引用 12 次
- TQL: Scaling Q-Functions with Transformers by Preventing Attention CollapsePerry Dong, Kuo-Han Hung, Alexander Swerdlow, Dorsa Sadigh 等ICML 2026 · 被引用 7 次
- DyWA: Dynamics-Adaptive World Action Model for Generalizable Non-Prehensile ManipulationJiangran Lyu, Ziming Li, Xuesong Shi, Chaoyi Xu 等ICCV 2025 · 被引用 2 次
它引用的顶会 Paper33
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee 等NeurIPS 2021 · 被引用 2,557 次
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 被引用 1,292 次
- Planning with Diffusion for Flexible Behavior SynthesisMichael Janner, Yilun Du, Joshua B. Tenenbaum, Sergey LevineICML 2022 · 被引用 1,115 次
相关 Paper
- Multi-Game Decision TransformersKuang-Huei Lee, Ofir Nachum, Mengjiao Yang, Lisa Lee 等NeurIPS 2022 · 被引用 279 次
- Meta-DT: Offline Meta-RL as Conditional Sequence Modeling with World Model DisentanglementZhi Wang, Li Zhang, Wenhao Wu, Yuanheng Zhu 等NeurIPS 2024 · 被引用 31 次
- Offline Trajectory Optimization for Offline Reinforcement LearningZiqi Zhao, Zhaochun Ren, Liu Yang, Yunsen Liang 等KDD 2025
- Diffusion Model is an Effective Planner and Data Synthesizer for Multi-Task Reinforcement LearningHaoran He, Chenjia Bai, Kang Xu, Zhuoran Yang 等NeurIPS 2023 · 被引用 165 次
- Transformer-based World Models Are Happy With 100k InteractionsJan Robine, Marc Höftmann, Tobias Uelwer, Stefan HarmelingICLR 2023 · 被引用 4 次
