Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation
Fanqing Meng, Jiaqi Liao, Xinyu Tan, Quanfeng Lu, Wenqi Shao, Kaipeng Zhang, Yu Cheng, Dianqi Li, Ping Luo
摘要
Text-to-video (T2V) models like Sora have made significant strides in visualizing complex prompts, which is increasingly viewed as a promising path towards constructing the universal world simulator. Cognitive psychologists believe that the foundation for achieving this goal is the ability to understand intuitive physics. However, the capacity of these models to accurately represent intuitive physics remains largely unexplored. To bridge this gap, we introduce PhyGen-Bench , a comprehensive Physics Generation Benchmark designed to evaluate physical commonsense correctness in T2V generation. PhyGenBench comprises 160 carefully crafted prompts across 27 distinct physical laws, spanning four fundamental domains, which could comprehensively assesses models' understanding of physical commonsense. Alongside PhyGenBench , we propose a novel evaluation framework called PhyGenEval . This framework employs a hierarchical evaluation structure utilizing appropriate advanced vision-language models and large language models to assess physical commonsense. Through Phy-GenBench and PhyGenEval , we can conduct large-scale automated assessments of T2V models' understanding of physical commonsense, which align closely with human feedback. Our evaluation results and in-depth analysis demonstrate that current models struggle to generate videos that comply with physical commonsense. Moreover, simply scaling up models or employing prompt engineering techniques is insufficient to fully address the challenges presented by Phy-GenBench (e.g., dynamic physical phenomenons). We hope this study will inspire the community to prioritize the learning of physical commonsense in these models beyond entertainment applications. We release the data and codes at https://github.com/OpenGVLab/PhyGenBench
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper49
- VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video GenerationHritik Bansal, Clark Peng, Yonatan Bitton, Roman Goldenberg 等ICLR 2026 · 被引用 146 次
- Genie Envisioner: A Unified World Foundation Platform for Robotic ManipulationYue Liao, Pengfei Zhou, Siyuan Huang, Donglin Yang 等ICLR 2026 · 被引用 136 次
- VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation ModelsXiangdong Zhang, Jiaqi Liao, Shaofeng Zhang, Fanqing Meng 等NeurIPS 2025 · 被引用 98 次
- WISA: World simulator assistant for physics-aware text-to-video generationJing Wang, Ao Ma, Ke Cao, Jun Zheng 等NeurIPS 2025 · 被引用 93 次
- NewtonGen: Physics-consistent and Controllable Text-to-Video Generation via Neural Newtonian DynamicsYu Yuan, Xijun Wang, Tharindu Wickremasinghe, Zeeshan Nadir 等ICLR 2026 · 被引用 46 次
它引用的顶会 Paper9
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras 等EMNLP 2021 · 被引用 937 次
- Vista: A Generalizable Driving World Model with High Fidelity and Versatile ControllabilityShenyuan Gao, Jiazhi Yang, Li Chen, Kashyap Chitta 等NeurIPS 2024 · 被引用 403 次
- Generating Long Videos of Dynamic ScenesTim Brooks, Janne Hellsten, Miika Aittala, Ting-Chun Wang 等NeurIPS 2022 · 被引用 152 次
- Evaluation of Text-to-Video Generation Models: A Dynamics PerspectiveMingxiang Liao, Hannan Lu, Qixiang Ye, Wangmeng Zuo 等NeurIPS 2024 · 被引用 89 次
相关 Paper
- PhyWorldBench: A Comprehensive Evaluation of Physical Realism in Text-to-Video ModelsJing Gu, Xian Liu, Yu Zeng, Ashwin Nagarajan 等ICLR 2026 · 被引用 29 次
- PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World UnderstandingWei Chow, Jiageng Mao, Boyi Li, Daniel Seita 等ICLR 2025 · 被引用 2 次
- VideoPhy: Evaluating Physical Commonsense for Video GenerationHritik Bansal, Zongyu Lin, Tianyi Xie, Zeshun Zong 等ICLR 2025 · 被引用 1 次
- PAI-Bench: A Comprehensive Benchmark For Physical AIFengzhe Zhou, Jiannan Huang, Jialuo Li, Deva Ramanan 等CVPR 2026 · 被引用 32 次
- Chain of Event-Centric Causal Thought for Physically Plausible Video GenerationZixuan Wang, Yixin Hu, Haolan Wang, Feng Chen 等CVPR 2026 · 被引用 8 次
