The Matrix: Infinite-Horizon World Generation with Real-Time Moving Control
Ruili Feng, Han Zhang, Zhilei Shu, Zhantao Yang, Longxiang Tang, Zhicai Wang, Andy Zheng, Jie Xiao, Zhiheng Liu, Ruihang Chu, Yukun Huang, Yu Liu, Hongyang Zhang
Abstract
We present The Matrix, a foundational realistic world simulator capable of generating infinitely long 720p high-fidelity real-scene video streams with real-time, responsive control in both first-and third-person perspectives. Trained on limited data from video games like Forza Horizon 5 and Cyberpunk 2077, complemented by large-scale unsupervised footage from real-world settings like Tokyo streets, The Matrix allows users to traverse diverse terrains-deserts, grasslands, water bodies, and urban landscapes-in continuous, uncut hour-long sequences. With speeds of up to 16 FPS, the system supports real-time interactivity and demonstrates zero-shot generalization, translating virtual game environments to real-world contexts where collecting continuous movement data is often infeasible. For example, The Matrix can simulate a BMW X3 driving through an office setting-an environment present in neither gaming data nor real-world sources. This approach showcases the potential of game data to advance robust world models, bridging the gap between simulations and real-world applications in scenarios with limited data. See https://github.com/MatrixTeam-AI/matrix, https://matrixteam-ai.github.io/pages/TheMatrix/ for code data and project page.
Figure 1: The Matrix is a foundational realistic world simulator capable of generating infinitely long 720p high-fidelity real-scene video streams with real-time, precise moving control. Due to size limitation, we recommend you to go to the website provided in the abstract for videos.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ad88c54b-00d7-47ac-a8ec-bbb25e4663fdCited by top-tier papers27
- LongLive: Real-time Interactive Long Video GenerationShuai Yang, Wei Huang, Ruihang Chu, Yicheng Xiao et al.ICLR 2026 · 241 citations
- WorldMem: Long-term Consistent World Simulation with MemoryZeqi Xiao, Yushi Lan, Yifan Zhou, Wenqi Ouyang et al.NeurIPS 2025 · 165 citations
- Video World Models with Long-term Spatial MemoryTong Wu, Shuai Yang, Ryan Po, Yinghao Xu et al.NeurIPS 2025 · 145 citations
- Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World ModelingHaoyu Wu, Diankun Wu, Tianyu He, Junliang Guo et al.ICLR 2026 · 89 citations
- Autoregressive Adversarial Post-Training for Real-Time Interactive Video GenerationShanchuan Lin, Ceyuan Yang, Hao He, Jianwen Jiang et al.NeurIPS 2025 · 89 citations
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Consistency ModelsYang Song, Prafulla Dhariwal, Mark Chen, Ilya SutskeverICML 2023 · 1,720 citations
Related papers
- Learning Interactive Real-World SimulatorsSherry Yang, Yilun Du, Seyed Kamyar Seyed Ghasemipour, Jonathan Tompson et al.ICLR 2024 · 399 citations
- DriveGAN: Towards a Controllable High-Quality Neural SimulationSeung Wook Kim, Jonah Philion, Antonio Torralba, Sanja FidlerCVPR 2021
- Genie: Generative Interactive EnvironmentsJake Bruce, Michael D. Dennis, Ashley Edwards, Jack Parker-Holder et al.ICML 2024 · 513 citations
- ReSim: Reliable World Simulation for Autonomous DrivingJiazhi Yang, Kashyap Chitta, Shenyuan Gao, Long Chen et al.NeurIPS 2025 · 53 citations
- Pre-Trained Video Generative Models as World SimulatorsHaoran He, Yang Zhang, Liang Lin, Zhongwen Xu et al.AAAI 2026 · 32 citations
