Endless World: Real-Time 3D-Aware Long Video Generation
Ke Zhang, Jiacong Xu, Yiqun Mei, Vishal M. Patel
Abstract
Producing long, coherent video sequences with stable 3D structure remains a major challenge, particularly in streaming scenarios. Motivated by this, we introduce Endless World, a real-time framework for infinite, 3D-consistent video generation.To support infinite video generation, we introduce a conditional autoregressive training strategy that aligns newly generated content with existing video frames. This design preserves long-range dependencies while remaining computationally efficient, enabling real-time inference on a single GPU without additional training overhead.Moreover, our Endless World integrates global 3D-aware attention to provide continuous geometric guidance across time. Our 3D injection mechanism enforces physical plausibility and geometric consistency throughout extended sequences, addressing key challenges in long-horizon and dynamic scene synthesis.Extensive experiments demonstrate that Endless World produces long, stable, and visually coherent videos, achieving competitive or superior performance to existing methods in both visual fidelity and spatial consistency. Our project has been available on https://bwgzk-keke.github.io/EndlessWorld/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4b8a931c-86f9-4a72-9787-07a464f09feeCited by top-tier papers1
Ask how each one uses itBuilds on24
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
- Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance FieldsJonathan T. Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman et al.ICCV 2021 · 2,700 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance FieldsJonathan T. Barron, Ben Mildenhall, Dor Verbin, Pratul P. Srinivasan et al.CVPR 2022 · 1,603 citations
Related papers
- SceneScape: Text-Driven Consistent Scene GenerationRafail Fridman, Amit Abecasis, Yoni Kasten, Tali DekelNeurIPS 2023 · 196 citations
- Infinite Nature: Perpetual View Generation of Natural Scenes from a Single ImageAndrew Liu, Ameesh Makadia, Richard Tucker, Noah Snavely et al.ICCV 2021 · 260 citations
- WorldReel: 4D Video Generation with Consistent Geometry and Motion ModelingShaoheng Fang, Hanwen Jiang, Yunpeng Bai, Niloy J. Mitra et al.CVPR 2026 · 3 citations
- LIVE: Long-horizon Interactive Video World ModelingJunchao Huang, Ziyang Ye, Xinting Hu, Tianyu He et al.ICML 2026 · 12 citations
- WorldGrow: Generating Infinite 3D WorldSikuang Li, Chen Yang, Jiemin Fang, Taoran Yi et al.AAAI 2026 · 10 citations
