The 3D-PC: a benchmark for visual perspective taking in humans and machines
Drew Linsley, Peisen Zhou, Alekh Karkada Ashok, Akash Nagaraj, Gaurav Gaonkar, Francis E. Lewis, Zygmunt Pizlo, Thomas Serre
摘要
Visual perspective taking (VPT) is the ability to perceive and reason about the perspectives of others. It is an essential feature of human intelligence, which develops over the first decade of life and requires an ability to process the 3D structure of visual scenes. A growing number of reports have indicated that deep neural networks (DNNs) become capable of analyzing 3D scenes after training on large image datasets. We investigated if this emergent ability for 3D analysis in DNNs is sufficient for VPT with the 3D perception challenge (3D-PC): a novel benchmark for 3D perception in humans and DNNs. The 3D-PC is comprised of three 3D-analysis tasks posed within natural scene images: 1. a simple test of object depth order, 2. a basic VPT task (VPT-basic), and 3. another version of VPT (VPT-Strategy) designed to limit the effectiveness of "shortcut" visual strategies. We tested human participants (N=33) and linearly probed or text-prompted over 300 DNNs on the challenge and found that nearly all of the DNNs approached or exceeded human accuracy in analyzing object depth order. Surprisingly, DNN accuracy on this task correlated with their object recognition performance. In contrast, there was an extraordinary gap between DNNs and humans on VPT-basic. Humans were nearly perfect, whereas most DNNs were near chance. Fine-tuning DNNs on VPT-basic brought them close to human performance, but they, unlike humans, dropped back to chance when tested on VPT-Strategy. Our challenge demonstrates that the training routines and architectures of today's DNNs are well-suited for learning basic 3D properties of scenes and objects but are ill-suited for reasoning about these properties as humans do. We release our 3D-PC datasets and code to help bridge this gap in 3D perception between humans and machines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- SpinBench: Perspective and Rotation as a Lens on Spatial Reasoning in VLMsYuyou Zhang, Radu Corcodel, Chiori Hori, Anoop Cherian 等ICLR 2026 · 被引用 11 次
- Perspective-Aware Reasoning in Vision-Language Models via Mental Imagery SimulationPhillip Y. Lee, Jihyeon Je, Chanho Park, Leonidas J. Guibas 等ICCV 2025 · 被引用 6 次
- Learning to See Through a Baby’s Eyes: Early Visual Diets Enable Robust Visual Intelligence in Humans and MachinesYusen Cai, Qing Lin, BHARGAVA SATYA NUNNA, Mengmi ZhangCVPR 2026 · 被引用 4 次
- Token Warping Helps MLLMs Look from Nearby ViewpointsPhillip Y. Lee, Chanho Park, Mingue Park, Seungwoo Yoo 等CVPR 2026 · 被引用 1 次
- Feat2GS: Probing Visual Foundation Models with Gaussian SplattingYue Chen, Xingyu Chen, Anpei Chen, Gerard Pons-Moll 等CVPR 2025
它引用的顶会 Paper36
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 被引用 3,632 次
相关 Paper
- SPARE3D: A Dataset for SPAtial REasoning on Three-View Line DrawingsWenyu Han, Siyuan Xiang, Chenhui Liu, Ruoyu Wang 等CVPR 2020
- 3DSRBENCH: A Comprehensive 3D Spatial Reasoning BenchmarkWufei Ma, Haoyu Chen, Guofeng Zhang, Yu-Cheng Chou 等ICCV 2025 · 被引用 15 次
- GePBench: Evaluating Fundamental Geometric Perception for Multimodal Large Language ModelsShangyu Xing, Changhao Xiang, Xinyu Liu, Zhangtai Wu 等ICML 2026
- Probing Neural Representations of Scene Perception in a Hippocampally Dependent Task Using Artificial Neural NetworksMarkus Frey, Christian F. Doeller, Caswell BarryCVPR 2023
- V-PROM: A Benchmark for Visual Reasoning Using Visual Progressive MatricesDamien Teney, Peng Wang, Jiewei Cao, Lingqiao Liu 等AAAI 2020 · 被引用 37 次
