A General Protocol to Probe Large Vision Models for 3D Physical Understanding
Guanqi Zhan, Chuanxia Zheng, Weidi Xie, Andrew Zisserman
摘要
Our objective in this paper is to probe large vision models to determine to what extent they 'understand' different physical properties of the 3D scene depicted in an image. To this end, we make the following contributions: (i) We introduce a general and lightweight protocol to evaluate whether features of an off-the-shelf large vision model encode a number of physical 'properties' of the 3D scene, by training discriminative classifiers on the features for these properties. The probes are applied on datasets of real images with annotations for the property. (ii) We apply this protocol to properties covering scene geometry, scene material, support relations, lighting, and view-dependent measures, and large vision models including CLIP, DINOv1, DINOv2, VQGAN, Stable Diffusion. (iii) We find that features from Stable Diffusion and DINOv2 are good for discriminative learning of a number of properties, including scene geometry, support relations, shadows and depth, but less performant for occlusion and material, while outperforming DINOv1, CLIP and VQGAN for all properties. (iv) It is observed that different time steps of Stable Diffusion features, as well as different transformer layers of DINO/CLIP/VQGAN, are good at different properties, unlocking potential applications of 3D physical understanding.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Into the Rabbit Hull: From Task-Relevant Concepts in DINO to Minkowski GeometryThomas Fel, Binxu Wang, Michael A. Lepori, Matthew Kowal 等ICLR 2026 · 被引用 28 次
- LikePhys: Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood PreferenceJianhao Yuan, Fabio Pizzati, Francesco Pinto, Lars Kunze 等ICLR 2026 · 被引用 18 次
- Visual Jenga: Discovering Object Dependencies via Counterfactual InpaintingAnand Bhattad, Konpat Preechakul, Alexei A. EfrosNeurIPS 2025 · 被引用 13 次
- How Much 3D Do Video Foundation Models Encode?Zixuan Huang, Xiang Li, Zhaoyang Lv, James M.CVPR 2026 · 被引用 10 次
- SeeThrough3D: Occlusion Aware 3D Control in Text-to-Image GenerationVaibhav Agrawal, Rishubh Parihar, Pradhaan Bhat, Ravi Kiran Sarvadevabhatla 等CVPR 2026 · 被引用 5 次
它引用的顶会 Paper30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao 等NeurIPS 2024 · 被引用 2,305 次
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu 等CVPR 2024 · 被引用 847 次
相关 Paper
- Cross-Modal and Uncertainty-Aware Agglomeration for Open-Vocabulary 3D Scene UnderstandingJinlong Li, Cristiano Saltori, Fabio Poiesi, Nicu SebeCVPR 2025
- Towards Spatially Consistent Image Generation: On Incorporating Intrinsic Scene Properties into Diffusion ModelsHyundo Lee, Suhyung Choi, Inwoo Hwang, Byoung-Tak ZhangAAAI 2026
- QUANTIPHY: A Quantitative Benchmark Evaluating Physical Reasoning Abilities of Vision-Language ModelsLi Puyin, Tiange Xiang, Ella Mao, Shirley Wei 等CVPR 2026 · 被引用 23 次
- View Transformation Robustness for Multi-View 3D Object Reconstruction with Reconstruction Error-Guided View SelectionQi Zhang, Zhouhang Luo, Tao Yu, Hui HuangAAAI 2025 · 被引用 1 次
- Localizing and Editing Knowledge In Text-to-Image Generative ModelsSamyadeep Basu, Nanxuan Zhao, Vlad I. Morariu, Soheil Feizi 等ICLR 2024 · 被引用 50 次
