Detecting Violations of Physical Common Sense in Images: A Challenge Dataset and Effective Model
Weibin Wu, Zitong Wang, Zhengjie Luo, Wenqing Chen, Zibin Zheng
Abstract
Vision-language models (VLMs) have achieved remarkable success in various vision-language tasks, such as image captioning and visual question answering. However, these models often lack physical common sense, frequently failing to identify visually evident violations of common physical principles. Therefore, evaluating the VLMs' understanding of physical common sense is essential, which has not yet been systematically explored in existing research. To fill this gap, we introduce PhyVIB (Physical Common Sense Violation Image Benchmark). This novel benchmark consists of 16,000 images across eight categories, aiming to systematically assess the VLMs' capability to detect violations of physical common sense in images. Our evaluations show that even the state-of-the-art VLMs perform poorly on PhyVIB, highlighting a significant area for improvement. In response, we propose PhyDetector, a two-stage fine-tuning framework to enhance the VLMs' capability to detect violations of physical common sense. The first stage involves supervised fine-tuning, which equips the VLM with essential concepts related to visual physical anomalies. The second stage utilizes group relative policy optimization to enhance the VLM's multimodal reasoning capability on physical plausibility. Experimental results show that the model fine-tuned with PhyDetector can significantly outperform the state-of-the-art VLMs in physical common sense understanding. Our artifacts are available at https://github.com/ZitongWang018/PhyVIB.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get b4fc156a-f6de-4478-8347-4c843f8d02bbCited by top-tier papers2
- Ref4D-VideoBench: Four-Dimensional Reference-Based Evaluation of Text-to-Video Generative ModelsJiajia Wei, YuJia He, Yuhan Hou, Hang Qi et al.CVPR 2026
- QRShield: Exploiting Vulnerabilities of Latent Diffusion Models for Preventing AI Art PlagiarismXunyue Mo, Weibin Wu, Qingrui Tu, Hang Wang et al.AAAI 2026
Related papers
- PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World UnderstandingWei Chow, Jiageng Mao, Boyi Li, Daniel Seita et al.ICLR 2025 · 2 citations
- VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video UnderstandingZongxia Li, Xiyang Wu, Guangyao Shi, Yubin Qin et al.NeurIPS 2025 · 38 citations
- DeepPhy: Benchmarking Agentic VLMs on Physical ReasoningXinrun Xu, Pi Bu, Ye Wang, Börje F. Karlsson et al.AAAI 2026 · 6 citations
- VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMsBrigitta Malagurski Törtei, Yasser Dahou, Ngoc Dung Huynh, Wamiq Reyaz Para et al.CVPR 2026 · 3 citations
- Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video GenerationFanqing Meng, Jiaqi Liao, Xinyu Tan, Quanfeng Lu et al.ICML 2025
