iVideoGPT: Interactive VideoGPTs are Scalable World Models
Jialong Wu, Shaofeng Yin, Ningya Feng, Xu He, Dong Li, Jianye Hao, Mingsheng Long
Abstract
World models empower model-based agents to interactively explore, reason, and plan within imagined environments for real-world decision-making. However, the high demand for interactivity poses challenges in harnessing recent advancements in video generative models for developing world models at scale. This work introduces Interactive VideoGPT (iVideoGPT), a scalable autoregressive transformer framework that integrates multimodal signals--visual observations, actions, and rewards--into a sequence of tokens, facilitating an interactive experience of agents via next-token prediction. iVideoGPT features a novel compressive tokenization technique that efficiently discretizes high-dimensional visual observations. Leveraging its scalable architecture, we are able to pre-train iVideoGPT on millions of human and robotic manipulation trajectories, establishing a versatile foundation that is adaptable to serve as interactive world models for a wide range of downstream tasks. These include action-conditioned video prediction, visual planning, and model-based reinforcement learning, where iVideoGPT achieves competitive performance compared with state-of-the-art methods. Our work advances the development of interactive general world models, bridging the gap between generative video models and practical model-based reinforcement learning applications. Code and pre-trained models are available at https://thuml.github.io/iVideoGPT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b5e5ae49-2ec3-407b-be6b-41df95c97f1cCited by top-tier papers34
- Ctrl-World: A Controllable Generative World Model for Robot ManipulationYanjiang Guo, Lucy Xiaoyang Shi, Jianyu Chen, Chelsea FinnICLR 2026 · 163 citations
- DreamDojo: A Real-Time Robot World Model from Large-Scale Human VideosShenyuan Gao, William Liang, Kaiyuan Zheng, Ayaan Malik et al.ICML 2026 · 96 citations
- Vid2World: Crafting Video Diffusion Models to Interactive World ModelsSiqiao Huang, Jialong Wu, Qixing Zhou, Shangchen Miao et al.ICLR 2026 · 68 citations
- RLVR-World: Training World Models with Reinforcement LearningJialong Wu, Shaofeng Yin, Ningya Feng, Mingsheng LongNeurIPS 2025 · 52 citations
- RoboScape: Physics-informed Embodied World ModelYu Shang, Xin Zhang, Yinzhou Tang, Lei Jin et al.NeurIPS 2025 · 43 citations
Builds on31
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
Related papers
- DrivingGPT: Unifying Driving World Modeling and Planning with Multi-Modal Autoregressive TransformersYuntao Chen, Yuqi Wang, Zhaoxiang ZhangICCV 2025 · 7 citations
- AdaWorld: Learning Adaptable World Models with Latent ActionsShenyuan Gao, Siyuan Zhou, Yilun Du, Jun Zhang et al.ICML 2025
- DriveGPT: Scaling Autoregressive Behavior Models for DrivingXin Huang, Eric M. Wolff, Paul Vernaza, Tung Phan-Minh et al.ICML 2025
- SAMPO: Scale-wise Autoregression with Motion Prompt for Generative World ModelsSen Wang, Jingyi Tian, Le Wang, Zhimin Liao et al.NeurIPS 2025 · 3 citations
- Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from VideosYi Chen, Yuying Ge, Weiliang Tang, Yizhuo Li et al.ICCV 2025 · 5 citations
