iVideoGPT: Interactive VideoGPTs are Scalable World Models
Jialong Wu, Shaofeng Yin, Ningya Feng, Xu He, Dong Li, Jianye Hao, Mingsheng Long
摘要
World models empower model-based agents to interactively explore, reason, and plan within imagined environments for real-world decision-making. However, the high demand for interactivity poses challenges in harnessing recent advancements in video generative models for developing world models at scale. This work introduces Interactive VideoGPT (iVideoGPT), a scalable autoregressive transformer framework that integrates multimodal signals--visual observations, actions, and rewards--into a sequence of tokens, facilitating an interactive experience of agents via next-token prediction. iVideoGPT features a novel compressive tokenization technique that efficiently discretizes high-dimensional visual observations. Leveraging its scalable architecture, we are able to pre-train iVideoGPT on millions of human and robotic manipulation trajectories, establishing a versatile foundation that is adaptable to serve as interactive world models for a wide range of downstream tasks. These include action-conditioned video prediction, visual planning, and model-based reinforcement learning, where iVideoGPT achieves competitive performance compared with state-of-the-art methods. Our work advances the development of interactive general world models, bridging the gap between generative video models and practical model-based reinforcement learning applications. Code and pre-trained models are available at https://thuml.github.io/iVideoGPT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper34
- Ctrl-World: A Controllable Generative World Model for Robot ManipulationYanjiang Guo, Lucy Xiaoyang Shi, Jianyu Chen, Chelsea FinnICLR 2026 · 被引用 163 次
- DreamDojo: A Real-Time Robot World Model from Large-Scale Human VideosShenyuan Gao, William Liang, Kaiyuan Zheng, Ayaan Malik 等ICML 2026 · 被引用 96 次
- Vid2World: Crafting Video Diffusion Models to Interactive World ModelsSiqiao Huang, Jialong Wu, Qixing Zhou, Shangchen Miao 等ICLR 2026 · 被引用 68 次
- RLVR-World: Training World Models with Reinforcement LearningJialong Wu, Shaofeng Yin, Ningya Feng, Mingsheng LongNeurIPS 2025 · 被引用 52 次
- RoboScape: Physics-informed Embodied World ModelYu Shang, Xin Zhang, Yinzhou Tang, Lei Jin 等NeurIPS 2025 · 被引用 43 次
它引用的顶会 Paper31
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
相关 Paper
- DrivingGPT: Unifying Driving World Modeling and Planning with Multi-Modal Autoregressive TransformersYuntao Chen, Yuqi Wang, Zhaoxiang ZhangICCV 2025 · 被引用 7 次
- AdaWorld: Learning Adaptable World Models with Latent ActionsShenyuan Gao, Siyuan Zhou, Yilun Du, Jun Zhang 等ICML 2025
- DriveGPT: Scaling Autoregressive Behavior Models for DrivingXin Huang, Eric M. Wolff, Paul Vernaza, Tung Phan-Minh 等ICML 2025
- SAMPO: Scale-wise Autoregression with Motion Prompt for Generative World ModelsSen Wang, Jingyi Tian, Le Wang, Zhimin Liao 等NeurIPS 2025 · 被引用 3 次
- Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from VideosYi Chen, Yuying Ge, Weiliang Tang, Yizhuo Li 等ICCV 2025 · 被引用 5 次
