Flow Equivariant World Models: Structured Memory for Dynamic Environments
Hansen Lillemark, Benhao Huang, Fangneng Zhan, Yilun Du, T. Anderson Keller
摘要
Embodied systems experience the world as 'a symphony of flows': a combination of many continuous streams of sensory input coupled to self-motion, interwoven with the dynamics of external objects. These sensory streams and the underlying dynamics of the world obey smooth, time-parameterized symmetries which existing world models ignore. Without a memory that respects this structure, partial observability presents a major obstacle to existing methods: each observation reveals only a fraction of the world, while unobserved regions continue to evolve. In this work, we introduce Flow Equivariant World Modeling, a framework that leverages time-parameterized symmetries within a latent memory for stable and accurate dynamics prediction over long horizons. The latent memory shifts and transforms equivariantly with self-motion and inferred external object motion, keeping information about out-of-view regions aligned as time progresses. We demonstrate the advantage of this framework over state-of-the-art diffusion, memory-augmented, and recurrent world model architectures on 2D and 3D partially observed video world modeling benchmarks. More broadly, our results suggest that predictive representations become more powerful when they are organized in line with the temporal and dynamical structure of the world they model. Project page: https://flowequivariantworldmodels.github.io/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper39
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 被引用 1,407 次
相关 Paper
- Learning 3D Persistent Embodied World ModelsSiyuan Zhou, Yilun Du, Yuncong Yang, Lei Han 等NeurIPS 2025 · 被引用 34 次
- Flow Equivariant Recurrent Neural NetworksAndy KellerNeurIPS 2025 · 被引用 8 次
- Future Dynamic 3D Reconstruction: A 3D World Model with Disentangled Ego-MotionNils Morbitzer, Jonathan Evers, Artem Savkin, Thomas Stauner 等ICML 2026
- WorldWeaver: Generating Long-Horizon Video Worlds via Rich PerceptionZhiheng Liu, Xueqing Deng, Shoufa Chen, Angtian Wang 等NeurIPS 2025 · 被引用 15 次
- Flow for Future: Geometric SE(3)-Equivariant Flow Matching for 3D Trajectory PredictionJunwei Wu, Yihang Liu, Ruixuan Yu, Jian SunICML 2026
