What Happens Next? Anticipating Future Motion by Generating Point Trajectories
Gabrijel Boduljak, Laurynas Karazija, Iro Laina, Christian Rupprecht, Andrea Vedaldi
摘要
We consider the problem of forecasting motion from a single image, i.e., predicting how objects in the world are likely to move, without the ability to observe other parameters such as the object velocities or the forces applied to them. We formulate this task as conditional generation of dense trajectory grids with a model that closely follows the architecture of modern video generators but outputs motion trajectories instead of pixels. This approach captures scene-wide dynamics and uncertainty, yielding more accurate and diverse predictions than prior regressors and generators. Although recent state-of-the-art video generators are often regarded as world models, we show that they struggle with forecasting motion from a single image, even in simple physical scenarios such as falling blocks or mechanical object interactions, despite fine-tuning on such data. We show that this limitation arises from the overhead of generating pixels rather than directly modeling motion.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Envisioning the Future, One Step at a TimeStefan Andreas Baumann, Jannik Wiese, Tommaso Martorella, M. Kalayeh 等CVPR 2026 · 被引用 4 次
- Learning Long-term Motion Embeddings for Efficient Kinematics GenerationNick Stracke, Kolja Bauer, Stefan Andreas Baumann, Miguel Ángel Bautista 等CVPR 2026 · 被引用 2 次
- VFMF: Dense Forecasting by Generating Foundation Model FeaturesGabrijel Boduljak, Yushi Lan, Christian Rupprecht, Andrea VedaldiICML 2026
- Generative Point Tracking and ForecastingXuanchen Lu, Ang Cao, Chao Feng, Andrew OwensCVPR 2026
它引用的顶会 Paper22
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
- Learning Interactive Real-World SimulatorsSherry Yang, Yilun Du, Seyed Kamyar Seyed Ghasemipour, Jonathan Tompson 等ICLR 2024 · 被引用 399 次
- TAPIR: Tracking Any Point with per-frame Initialization and temporal RefinementCarl Doersch, Yi Yang, Mel Vecerík, Dilara Gokay 等ICCV 2023 · 被引用 297 次
相关 Paper
- Compositional Video PredictionYufei Ye, Maneesh Singh, Abhinav Gupta, Shubham TulsianiICCV 2019 · 被引用 84 次
- Learn the Force We Can: Enabling Sparse Motion Control in Multi-Object Video GenerationAram Davtyan, Paolo FavaroAAAI 2024 · 被引用 7 次
- InterDyn: Controllable Interactive Dynamics with Video Diffusion ModelsRick Akkerman, Haiwen Feng, Michael J. Black, Dimitrios Tzionas 等CVPR 2025
- Building 3D Representations and Generating Motions From a Single Image via Video-GenerationWeiming Zhi, Ziyong Ma, Tianyi Zhang, Matthew Johnson-RobersonNeurIPS 2025 · 被引用 1 次
- Motion Modes: What Could Happen Next?Karran Pandey, Yannick Hold-Geoffroy, Matheus Gadelha, Niloy J. Mitra 等CVPR 2025
