PhyWorldBench: A Comprehensive Evaluation of Physical Realism in Text-to-Video Models
Jing Gu, Xian Liu, Yu Zeng, Ashwin Nagarajan, Fangrui Zhu, Daniel Hong, Yue Fan, Qianqi Yan, Kaiwen Zhou, Ming-Yu Liu, Xin Wang
Abstract
Figure 2 : Success rates of video generation models on PhyWorldBench. Among open-source models, Wanx demonstrated the highest performance, while Pika achieved the best results among proprietary models with a success rate of 0.262. Despite these advancements, substantial progress remains necessary to refine the capability of these models to accurately simulate the intricate dynamics of the real world. ical phenomena with varying prompt types, deriving targeted recommendations for crafting prompts that enhance fidelity to physical principles. INTRODUCTION The field of video generation has made remarkable progress, with models producing visually compelling and often photorealistic outputs. These advances have enabled transformative applications across industries such as entertainment, education, and scientific visualization. However, despite their visual fidelity, do video generation models truly understand the laws of physics in the real world? To answer this question, we introduce PhyWorldBench, a rigorous benchmark designed to evaluate how well video generation models can simulate real-world physics. As illustrated in Figure 1 , PhyWorldBench systematically tests models across multiple levels of physical phenomena, from fundamental concepts like object motion to complex dynamics, including rigid body interactions and human/animal motion. Additionally, we propose a novel Anti-Physics category, where prompts deliberately violate real-world physics. On one hand, this design verifies whether models genuinely understand physical laws-rather than merely reproducing patterns from real-world training data. On the other hand, anti-physics content itself holds practical value in creative applications, where imaginative or otherwise impossible scenarios are beneficial. We meticulously designed and annotated 1,050 prompts and the standard set for each prompt individually to cover a broad range of physical scenarios. This substantial annotation work ensures that our benchmark is both comprehensive and precise, allowing for a more thorough assessment of video generation models' capabilities. Furthermore, we present a context-aware-prompt metric using MLLM (OpenAI Team, 2024; Gemini Team, 2024) , which directly assesses if the video satisfies the physics standards or not. Such evaluation not only provided an unbiased metric but also significantly reduced the evaluation cost. To examine the current status of video generation models and provide a detailed analysis, we selected five proprietary models-
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ad0df987-9ec6-4639-8e93-9a7ec5a9ead6Cited by top-tier papers5
- 4DWorldBench: A Comprehensive Evaluation Framework for 3D/4D World Generation ModelsYiting Lu, Wei Luo, Peiyan Tu, Haoran Li et al.CVPR 2026 · 10 citations
- PhysInOne: Visual Physics Learning and Reasoning in One SuiteSiyuan Zhou, Hejun Wang, Hu Cheng, Jinxi Li et al.CVPR 2026 · 9 citations
- Struct2D: A Perception-Guided Framework for Spatial Reasoning in MLLMsFangrui Zhu, Hanhui Wang, Yiming Xie, Jing Gu et al.NeurIPS 2025 · 7 citations
- SeeU: Seeing the Unseen World via 4D Dynamics-aware GenerationYu Yuan, Tharindu Wickremasinghe, Zeeshan Nadir, Xijun Wang et al.CVPR 2026 · 3 citations
- Lighting-grounded Video Generation with Renderer-based Agent ReasoningZiqi Cai, Taoyu Yang, Zheng Chang, Si Li et al.CVPR 2026 · 3 citations
Builds on8
- VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video GenerationHritik Bansal, Clark Peng, Yonatan Bitton, Roman Goldenberg et al.ICLR 2026 · 146 citations
- CogVideo: Large-scale Pretraining for Text-to-Video Generation via TransformersWenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu et al.ICLR 2023 · 116 citations
- Evaluation of Text-to-Video Generation Models: A Dynamics PerspectiveMingxiang Liao, Hannan Lu, Qixiang Ye, Wangmeng Zuo et al.NeurIPS 2024 · 89 citations
- VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video GenerationXuan He, Dongfu Jiang, Ge Zhang, Max Ku et al.EMNLP 2024 · 20 citations
- VideoPhy: Evaluating Physical Commonsense for Video GenerationHritik Bansal, Zongyu Lin, Tianyi Xie, Zeshun Zong et al.ICLR 2025 · 1 citation
Related papers
- Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video GenerationFanqing Meng, Jiaqi Liao, Xinyu Tan, Quanfeng Lu et al.ICML 2025
- Impossible VideosZechen Bai, Hai Ci, Mike Zheng ShouICML 2025
- PAI-Bench: A Comprehensive Benchmark For Physical AIFengzhe Zhou, Jiannan Huang, Jialuo Li, Deva Ramanan et al.CVPR 2026 · 32 citations
- Evaluating Newtonian Mechanics in Video Generative Models with Real Physical SystemsAntonios Tragoudaras, Chenyu Zhang, Daniil Cherniavskii, Antonis Vozikis et al.ICML 2026 · 39 citations
- PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World UnderstandingWei Chow, Jiageng Mao, Boyi Li, Daniel Seita et al.ICLR 2025 · 2 citations
