VMem: Consistent Interactive Video Scene Generation with Surfel-Indexed View Memory
Runjia Li, Philip Torr, Andrea Vedaldi, Tomas Jakab
摘要
We propose a novel memory module for building video generators capable of interactively exploring environments. Previous approaches have achieved similar results either by out-painting 2D views of a scene while incrementally reconstructing its 3D geometry—which quickly accumulates errors—or by using video generators with a short context window, which struggle to maintain scene coherence over the long term. To address these limitations, we introduce Surfel-Indexed View Memory (VMem), a memory module that remembers past views by indexing them geometrically based on the 3D surface elements (surfels) they have observed. VMem enables efficient retrieval of the most relevant past views when generating new ones. By focusing only on these relevant views, our method produces consistent explorations of imagined environments at a fraction of the computational cost required to use all past views as context. We evaluate our approach on challenging longterm scene synthesis benchmarks and demonstrate superior performance compared to existing methods in maintaining scene coherence and camera control.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World ModelingWenqiang Sun, Haiyu Zhang, Haoyuan Wang, Junta Wu 等ICML 2026 · 被引用 108 次
- Mixture of Contexts for Long Video GenerationShengqu Cai, Ceyuan Yang, Lvmin Zhang, Yuwei Guo 等ICLR 2026 · 被引用 92 次
- World-In-World: World Models in a Closed-Loop WorldJiahan Zhang, Muqing Jiang, Nanru Dai, Taiming Lu 等ICLR 2026 · 被引用 46 次
- Spatia: Video Generation with Updatable Spatial MemoryJinjing Zhao, Fangyun Wei, Zhening Liu, Hongyang Zhang 等CVPR 2026 · 被引用 37 次
- VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric ControlSixiao Zheng, Minghao Yin, Wenbo Hu, Xiaoyu Li 等CVPR 2026 · 被引用 27 次
它引用的顶会 Paper30
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- MVSNeRF: Fast Generalizable Radiance Field Reconstruction from Multi-View StereoAnpei Chen, Zexiang Xu, Fuqiang Zhao, Xiaoshuai Zhang 等ICCV 2021 · 被引用 1,024 次
- LRM: Large Reconstruction Model for Single Image to 3DYicong Hong, Kai Zhang, Jiuxiang Gu, Sai Bi 等ICLR 2024 · 被引用 813 次
- RegNeRF: Regularizing Neural Radiance Fields for View Synthesis from Sparse InputsMichael Niemeyer, Jonathan T. Barron, Ben Mildenhall, Mehdi S. M. Sajjadi 等CVPR 2022 · 被引用 513 次
- Genie: Generative Interactive EnvironmentsJake Bruce, Michael D. Dennis, Ashley Edwards, Jack Parker-Holder 等ICML 2024 · 被引用 513 次
相关 Paper
- 3D Scene Prompting for Scene-Consistent Camera-Controllable Video GenerationJoungBin Lee, Jaewoo Jung, Jisang Han, Takuya Narihira 等ICLR 2026 · 被引用 14 次
- Video World Models with Long-term Spatial MemoryTong Wu, Shuai Yang, Ryan Po, Yinghao Xu 等NeurIPS 2025 · 被引用 145 次
- Scenepainter: Semantically Consistent Perpetual 3D Scene Generation with Concept Relation AlignmentChong Xia, Shengjun Zhang, Fangfu Liu, Chang Liu 等ICCV 2025 · 被引用 2 次
- Geometry-as-context: Modulating Explicit 3D in Scene-consistent Video Generation to Geometry ContextJiaKui Hu, Jialun Liu, Liying Yang, Xinliang Zhang 等CVPR 2026 · 被引用 7 次
- Scalable Spatial Memory for Scene Rendering and NavigationWen-Cheng Chen, Chu-Song Chen, Wei-Chen Chiu, Min-Chun HuAAAI 2023 · 被引用 1 次
