VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation
Hritik Bansal, Clark Peng, Yonatan Bitton, Roman Goldenberg, Aditya Grover, Kai-Wei Chang
摘要
Large-scale video generative models, capable of creating realistic videos of diverse visual concepts, are strong candidates for general-purpose physical world simulators. However, their adherence to physical commonsense across real-world actions remains unclear (e.g., playing tennis, backflip). Existing benchmarks suffer from limitations such as limited size, lack of human evaluation, sim-to-real gaps, and absence of fine-grained physical rule analysis. To address this, we introduce VideoPhy-2, an action-centric dataset for evaluating physical commonsense in generated videos. We curate 4000 diverse and detailed prompts for video synthesis from modern generative models. We perform human evaluation that assesses semantic adherence, physical commonsense, and grounding of physical rules in the generated videos. Our findings reveal major shortcomings, with even the best model achieving only joint performance (i.e., high semantic and physical commonsense adherence) on the hard subset of VideoPhy-2. We find that the models particularly struggle with conservation laws like mass and momentum. Finally, we also train VideoPhy-2-AutoEval, an automatic evaluator for fast, reliable assessment on our dataset. Overall, VideoPhy-2 serves as a rigorous benchmark, exposing critical gaps in video generative models and guiding future research in physically-grounded video generation. The data and code is available at https://videophy2.github.io/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation ModelsXiangdong Zhang, Jiaqi Liao, Shaofeng Zhang, Fanqing Meng 等NeurIPS 2025 · 被引用 98 次
- PhysCtrl: Generative Physics for Controllable and Physics-Grounded Video GenerationChen Wang, Chuhao Chen, Yiming Huang, Zhiyang Dou 等NeurIPS 2025 · 被引用 50 次
- NewtonGen: Physics-consistent and Controllable Text-to-Video Generation via Neural Newtonian DynamicsYu Yuan, Xijun Wang, Tharindu Wickremasinghe, Zeeshan Nadir 等ICLR 2026 · 被引用 46 次
- Evaluating Newtonian Mechanics in Video Generative Models with Real Physical SystemsAntonios Tragoudaras, Chenyu Zhang, Daniil Cherniavskii, Antonis Vozikis 等ICML 2026 · 被引用 39 次
- PAI-Bench: A Comprehensive Benchmark For Physical AIFengzhe Zhou, Jiannan Huang, Jialuo Li, Deva Ramanan 等CVPR 2026 · 被引用 32 次
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video GeneratorsLevon Khachatryan, Andranik Movsisyan, Vahram Tadevosyan, Roberto Henschel 等ICCV 2023 · 被引用 800 次
相关 Paper
- Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video GenerationFanqing Meng, Jiaqi Liao, Xinyu Tan, Quanfeng Lu 等ICML 2025
- VideoPhy: Evaluating Physical Commonsense for Video GenerationHritik Bansal, Zongyu Lin, Tianyi Xie, Zeshun Zong 等ICLR 2025 · 被引用 1 次
- PhyWorldBench: A Comprehensive Evaluation of Physical Realism in Text-to-Video ModelsJing Gu, Xian Liu, Yu Zeng, Ashwin Nagarajan 等ICLR 2026 · 被引用 29 次
- PhysVid: Physics Aware Local Conditioning for Generative Video ModelsSaurabh Pathak, Elahe Arani, Mykola Pechenizkiy, Bahram ZonoozCVPR 2026 · 被引用 6 次
- Rethinking Video Generation Model for the Embodied WorldYufan Deng, Zilin Pan, Hongyu Zhang, Xiaojie Li 等ICML 2026 · 被引用 24 次
