Yume1.5: A Text-Controlled Interactive World Generation Model
Xiaofeng Mao, Zhen Li, Chuanhao Li, Xiaojie Xu, Kaining Ying, Kaipeng Zhang
Abstract
Recent approaches have demonstrated the promise of using diffusion models to generate interactive and explorable worlds. However, most of these methods face critical challenges such as excessively large parameter sizes, reliance on lengthy inference steps, and rapidly growing historical context, which severely limit real-time performance and lack text-controlled generation capabilities.To address these challenges, we propose Yume1.5, a novel framework designed to generate realistic, interactive, and continuous worlds from a single image or text prompt. Yume1.5 achieves this through a carefully designed framework that supports keyboard-based exploration of the generated worlds. The framework comprises three core components: (1) a long-video generation method combining unified context compression and linear attention; (2) a context compression-based bidirectional attention distillation approach with an enhanced text embedding scheme for real-time streaming video generation. Yume1.5 achieves an average generation speed of 12 fps at 540p resolution using only a single A100 GPU; (3) a text-controlled method for generating world events. We have provided the codebase in the supplementary material. The model weights and full codebase will be made public.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ce509c5b-aaf8-4547-a67f-32bf381fe3b7Cited by top-tier papers3
- WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World ModelingWenqiang Sun, Haiyu Zhang, Haoyuan Wang, Junta Wu et al.ICML 2026 · 108 citations
- Infinite-World: Scaling Interactive World Models to 1000-Frame Horizons via Pose-Free Hierarchical MemoryRuiqi Wu, Xuanhua He, Meng Cheng, Tianyu Yang et al.ICML 2026 · 29 citations
- Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video GenerationHongzhou Zhu, Min Zhao, Guande He, Hang Su et al.ICML 2026
Builds on20
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Diffusion Forcing: Next-token Prediction Meets Full-Sequence DiffusionBoyuan Chen, Diego Marti Monso, Yilun Du, Max Simchowitz et al.NeurIPS 2024 · 751 citations
- Improved Distribution Matching Distillation for Fast Image SynthesisTianwei Yin, Michaël Gharbi, Taesung Park, Richard Zhang et al.NeurIPS 2024 · 728 citations
- Self Forcing: Bridging the Train-Test Gap in Autoregressive Video DiffusionXun Huang, Zhengqi Li, Guande He, Mingyuan Zhou et al.NeurIPS 2025 · 628 citations
Related papers
- MotionStream: Real-Time Video Generation with Interactive Motion ControlsJoonghyuk Shin, Zhengqi Li, Richard Zhang, Jun-Yan Zhu et al.ICLR 2026 · 79 citations
- StreamDiT: Real-Time Streaming Text-to-Video GenerationAkio Kodaira, Tingbo Hou, Ji Hou, Markos Georgopoulos et al.CVPR 2026 · 45 citations
- LongLive: Real-time Interactive Long Video GenerationShuai Yang, Wei Huang, Ruihang Chu, Yicheng Xiao et al.ICLR 2026 · 241 citations
- FloodDiffusion: Tailored Diffusion Forcing for Streaming Motion GenerationYIYI CAI, Yuhan Wu, Kunhang Li, YOU ZHOU et al.CVPR 2026 · 14 citations
- UCM: Unified Modeling of Camera Control and Memory with Time-aware Positional Encoding Warping for World ModelsTianxing Xu, Zi-Xuan Wang, Guangyuan Wang, Li Hu et al.SIGGRAPH 2026
