Envisioning the Future, One Step at a Time
Stefan Andreas Baumann, Jannik Wiese, Tommaso Martorella, M. Kalayeh, Björn Ommer
Abstract
Accurately anticipating how complex, open-world scenes will evolve requires models that represent uncertainty, simulate along extended interaction chains, and efficiently explore many plausible futures. Yet most existing approaches rely on dense video or latent-space prediction, expending substantial capacity on dense appearance rather than on the underlying sparse trajectories of points in the scene. This makes large-scale exploration of future hypotheses costly and limits performance when long-horizon, multi-modal motion is essential. We address this by formulating the prediction of open-set future scene dynamics as step-wise inference over sparse point trajectories. Our autoregressive diffusion model advances these trajectories through short, locally predictable transitions, explicitly modeling the growth of uncertainty over time. This dynamics-centric representation enables fast rollout of thousands of diverse futures from a single image, optionally guided by initial constraints on motion, while maintaining physical plausibility and long-range coherence.We further introduce OWM, a benchmark for open-world motion prediction based on diverse in-the-wild videos, to evaluate accuracy and variability of predicted trajectory distributions under real-world uncertainty. Our method matches or surpasses dense simulators in predictive accuracy while achieving orders-of-magnitude higher sampling speed, making open-world future prediction both scalable and practical.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on40
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
Related papers
- What Happens Next? Anticipating Future Motion by Generating Point TrajectoriesGabrijel Boduljak, Laurynas Karazija, Iro Laina, Christian Rupprecht et al.ICLR 2026 · 10 citations
- Autoscape: Geometry-Consistent Long-Horizon Scene GenerationJiacheng Chen, Ziyu Jiang, Mingfu Liang, Bingbing Zhuang et al.ICCV 2025
- MoMaps: Semantics-Aware Scene Motion Generation with Motion MapsJiahui Lei, Kyle Genova, George Kopanas, Noah Snavely et al.ICCV 2025 · 1 citation
- DriveX: Omni Scene Modeling for Learning Generalizable World Knowledge in Autonomous DrivingChen Shi, Shaoshuai Shi, Kehua Sheng, Bo Zhang et al.ICCV 2025
- ExtDM: Distribution Extrapolation Diffusion Model for Video PredictionZhicheng Zhang, Junyao Hu, Wentao Cheng, Danda Pani Paudel et al.CVPR 2024 · 24 citations
