Evaluating Newtonian Mechanics in Video Generative Models with Real Physical Systems
Antonios Tragoudaras, Chenyu Zhang, Daniil Cherniavskii, Antonis Vozikis, Thijmen Nijdam, Derck Prinzhorn, Márk Bodrácska, Nicu Sebe, Andrii Zadaianchuk, Efstratios Gavves
摘要
Recent advances in image and video generation raise hopes that these models possess world modeling capabilities—the ability to generate realistic, physically plausible videos. This could revolutionize applications in robotics, autonomous driving, and scientific simulation. However, before treating these models as world models, we must ask: Do they adhere to physical laws? Current evaluation methods rely on subjective judgments or trajectory matching, limiting their usage for physical reasoning estimation, where many generations could be physically plausible. Thus, we introduce Morpheus , one of the first physics-informed evaluation frameworks for measuring the ability of video generation models to comprehend Newtonian dynamics. Morpheus features 130 real-world videos capturing physical phenomena, guided by conservation laws. Using those as conditioning for video generation, we assess physical plausibility leveraging interpretable metrics evaluated with respect to infallible conservation laws known per physical setting, leveraging advances in physics-informed neural networks and vision-language foundation models. Importantly, Morpheus targets controlled Newtonian rigid-body settings to enable quantitative checks. Our findings reveal that even with advanced prompting and video conditioning, contemporary models struggle to encode physical principles despite generating aesthetically pleasing videos. Code and data available here .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- PhysInOne: Visual Physics Learning and Reasoning in One SuiteSiyuan Zhou, Hejun Wang, Hu Cheng, Jinxi Li 等CVPR 2026 · 被引用 9 次
- SVBench: Evaluation of Video Generation Models on Social ReasoningWenshuo Peng, Gongxuan Wang, Tianmeng Yang, Chuanhao Li 等CVPR 2026 · 被引用 5 次
- SeeU: Seeing the Unseen World via 4D Dynamics-aware GenerationYu Yuan, Tharindu Wickremasinghe, Zeeshan Nadir, Xijun Wang 等CVPR 2026 · 被引用 3 次
- Benchmarking Single-Factor Physical Video-to-Audio GenerationTingle Li, Siddharth Gururani, Kevin Shih, Gantavya Bhatt 等CVPR 2026 · 被引用 2 次
它引用的顶会 Paper22
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao 等NeurIPS 2024 · 被引用 2,305 次
- Depth Anything 3: Recovering the Visual Space from Any ViewsHaotong Lin, Sili Chen, Jun Hao Liew, Donny Y. Chen 等ICLR 2026 · 被引用 720 次
- Learning Interactive Real-World SimulatorsSherry Yang, Yilun Du, Seyed Kamyar Seyed Ghasemipour, Jonathan Tompson 等ICLR 2024 · 被引用 399 次
相关 Paper
- PhyWorldBench: A Comprehensive Evaluation of Physical Realism in Text-to-Video ModelsJing Gu, Xian Liu, Yu Zeng, Ashwin Nagarajan 等ICLR 2026 · 被引用 29 次
- Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video GenerationFanqing Meng, Jiaqi Liao, Xinyu Tan, Quanfeng Lu 等ICML 2025
- VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video GenerationHritik Bansal, Clark Peng, Yonatan Bitton, Roman Goldenberg 等ICLR 2026 · 被引用 146 次
- MoReGen: Multi-Agent Motion-Reasoning Engine for Code-based Text-to-Video SynthesisXiangyu Bai, He Liang, Bishoy Galoaa, Utsav Nandi 等CVPR 2026 · 被引用 6 次
- PHANTOM: Physics-Infused Video Generation via Joint Modeling of Visual and Latent Physical DynamicsYing Shen, Jerry Xiong, Tianjiao Yu, Ismini LourentzouCVPR 2026 · 被引用 12 次
