Lune

CVPR2026顶会

VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object Detection

Yang Cao, Feize Wu, Dave Chen, Yingji Zhong, Lanqing Hong, Dan Xu

2026年份
6被引次数

摘要

Current multi-view indoor 3D object detectors rely on sensor geometry that is costly to obtain-i.e., precisely calibrated multi-view camera poses-to fuse multi-view information into a global scene representation, limiting deployment in real-world scenes. We target a more practical setting: Sensor-Geometry-Free (SG-Free) multi-view indoor 3D object detection, where there are no sensor-provided geometric inputs (multi-view poses or depth). Recent Visual Geometry Grounded Transformer (VGGT) shows that strong 3D cues can be inferred directly from images. Building on this insight, we present VGGT-Det, the first framework tailored for SG-Free multi-view indoor 3D object detection. Rather than merely consuming VGGT predictions, our method integrates VGGT encoder into a transformer-based pipeline. To effectively leverage both the semantic and geometric priors from inside VGGT, we introduce two novel key components: (i) Attention-Guided Query Generation (AG): exploits VGGT attention maps as semantic priors to initialize object queries, improving localization by focusing on object regions while preserving global spatial structure; (ii) Query-Driven Feature Aggregation (QD): a learnable See-Query interacts with object queries to 'see' what they need, and then dynamically aggregates multi-level geometric features across VGGT layers that progressively lift 2D features into 3D. Experiments show that VGGT-Det significantly surpasses the best-performing method in the SG-Free setting by 4.4 and 8.6 mAP@0.25 on ScanNet and ARKitScenes, respectively. Ablation study shows that VGGT's internally learned semantic and geometric priors can be effectively leveraged by our AG and QD. Source code and pre-trained models are available at the GitHub project page.

Current methods [4, 32, 36, 46, 47] predominantly rely on sensor-derived geometric inputs, i.e., precisely calibrated multi-view camera poses or depth. Although these methods achieve strong performance, their reliance on expensive and often inaccessible sensor-derived geometric inputs [3, 5] severely limits scalability and real-world deployment.

We instead consider a more practical setting: performing indoor 3D detection from multi-view images without sensor-derived geometric inputs. We refer to this as Sensor-Geometry-Free (SG-Free) multi-view indoor 3D object detection. This setting is highly challenging because it eliminates the sensor-derived multi-view camera poses and depth. Recent advances in feed-forward 3D reconstruction have shown that 3D structure can be inferred directly from unposed 2D images [15,37,41,43,49,50]. These develop-

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext d82b1f00-e2ab-442d-a36a-53bf7879c91b

它引用的顶会 Paper25

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖