DeepPhy: Benchmarking Agentic VLMs on Physical Reasoning
Xinrun Xu, Pi Bu, Ye Wang, Börje F. Karlsson, Ziming Wang, Tengtao Song, Qi Zhu, Jun Song, Zhiming Ding, Bo Zheng
摘要
Although Vision Language Models (VLMs) exhibit strong perceptual abilities and impressive visual reasoning, they struggle with attention to detail and precise action planning in complex, dynamic environments, leading to subpar performance. Real-world tasks typically require complex interactions, advanced spatial reasoning, long-term planning, and continuous strategy refinement, usually necessitating understanding the physics rules of the target scenario. However, evaluating these capabilities in real-world scenarios is often prohibitively expensive. To bridge this gap, we introduce DeepPHY, a novel benchmark framework designed to systematically evaluate VLMs' understanding and reasoning about fundamental physical principles through a series of challenging simulated environments. DeepPHY integrates multiple physical reasoning environments of varying difficulty levels and incorporates fine-grained evaluation metrics. Our evaluation finds that even state-of-the-art VLMs struggle to translate descriptive physical knowledge into precise, predictive control.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper8
- iVideoGPT: Interactive VideoGPTs are Scalable World ModelsJialong Wu, Shaofeng Yin, Ningya Feng, Xu He 等NeurIPS 2024 · 被引用 177 次
- CivRealm: A Learning and Reasoning Odyssey in Civilization for Decision-Making AgentsSiyuan Qi, Shuo Chen, Yexin Li, Xiangyu Kong 等ICLR 2024 · 被引用 35 次
- I-PHYRE: Interactive Physical ReasoningShiqian Li, Kewen Wu, Chi Zhang, Yixin ZhuICLR 2024 · 被引用 16 次
- PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World UnderstandingWei Chow, Jiageng Mao, Boyi Li, Daniel Seita 等ICLR 2025 · 被引用 2 次
- AndroidWorld: A Dynamic Benchmarking Environment for Autonomous AgentsChristopher Rawles, Sarah Clinckemaillie, Yifan Chang, Jonathan Waltz 等ICLR 2025
相关 Paper
- SIMPACT: Simulation-Enabled Action Planning using Vision-Language ModelsHaowen Liu, Shaoxiong Yao, Haonan Chen, Jiawei Gao 等CVPR 2026 · 被引用 8 次
- Detecting Violations of Physical Common Sense in Images: A Challenge Dataset and Effective ModelWeibin Wu, Zitong Wang, Zhengjie Luo, Wenqing Chen 等ACM MM 2025
- PAI-Bench: A Comprehensive Benchmark For Physical AIFengzhe Zhou, Jiannan Huang, Jialuo Li, Deva Ramanan 等CVPR 2026 · 被引用 32 次
- Beyond Static Vision: Scene Dynamic Field Unlocks Intuitive Physics Understanding in Multi-modal Large Language ModelsNanxi Li, Xiang Wang, Yuanjie Chen, Haode Zhang 等ICLR 2026
- Evaluating Vision-Language Models as Evaluators in Path PlanningMohamed Aghzal, Xiang Yue, Erion Plaku, Ziyu YaoCVPR 2025
