VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object Detection
Yang Cao, Feize Wu, Dave Chen, Yingji Zhong, Lanqing Hong, Dan Xu
Abstract
Current multi-view indoor 3D object detectors rely on sensor geometry that is costly to obtain-i.e., precisely calibrated multi-view camera poses-to fuse multi-view information into a global scene representation, limiting deployment in real-world scenes. We target a more practical setting: Sensor-Geometry-Free (SG-Free) multi-view indoor 3D object detection, where there are no sensor-provided geometric inputs (multi-view poses or depth). Recent Visual Geometry Grounded Transformer (VGGT) shows that strong 3D cues can be inferred directly from images. Building on this insight, we present VGGT-Det, the first framework tailored for SG-Free multi-view indoor 3D object detection. Rather than merely consuming VGGT predictions, our method integrates VGGT encoder into a transformer-based pipeline. To effectively leverage both the semantic and geometric priors from inside VGGT, we introduce two novel key components: (i) Attention-Guided Query Generation (AG): exploits VGGT attention maps as semantic priors to initialize object queries, improving localization by focusing on object regions while preserving global spatial structure; (ii) Query-Driven Feature Aggregation (QD): a learnable See-Query interacts with object queries to 'see' what they need, and then dynamically aggregates multi-level geometric features across VGGT layers that progressively lift 2D features into 3D. Experiments show that VGGT-Det significantly surpasses the best-performing method in the SG-Free setting by 4.4 and 8.6 mAP@0.25 on ScanNet and ARKitScenes, respectively. Ablation study shows that VGGT's internally learned semantic and geometric priors can be effectively leveraged by our AG and QD. Source code and pre-trained models are available at the GitHub project page.
Current methods [4, 32, 36, 46, 47] predominantly rely on sensor-derived geometric inputs, i.e., precisely calibrated multi-view camera poses or depth. Although these methods achieve strong performance, their reliance on expensive and often inaccessible sensor-derived geometric inputs [3, 5] severely limits scalability and real-world deployment.
We instead consider a more practical setting: performing indoor 3D detection from multi-view images without sensor-derived geometric inputs. We refer to this as Sensor-Geometry-Free (SG-Free) multi-view indoor 3D object detection. This setting is highly challenging because it eliminates the sensor-derived multi-view camera poses and depth. Recent advances in feed-forward 3D reconstruction have shown that 3D structure can be inferred directly from unposed 2D images [15,37,41,43,49,50]. These develop-
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d82b1f00-e2ab-442d-a36a-53bf7879c91bBuilds on25
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- BEVDepth: Acquisition of Reliable Depth for Multi-View 3D Object DetectionYinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang et al.AAAI 2023 · 954 citations
- LightGlue: Local Feature Matching at Light SpeedPhilipp Lindenberger, Paul-Edouard Sarlin, Marc PollefeysICCV 2023 · 936 citations
- An End-to-End Transformer Model for 3D Object DetectionIshan Misra, Rohit Girdhar, Armand JoulinICCV 2021 · 602 citations
- PETRv2: A Unified Framework for 3D Perception from Multi-Camera ImagesYingfei Liu, Junjie Yan, Fan Jia, Shuailin Li et al.ICCV 2023 · 513 citations
Related papers
- GGPT: Geometry-Grounded Point TransformerYutong Chen, Yiming Wang, Xucong Zhang, Sergey Prokudin et al.CVPR 2026 · 2 citations
- Viewpoint Equivariance for Multi-View 3D Object DetectionDian Chen, Jie Li, Vitor Guizilini, Rares Ambrus et al.CVPR 2023
- Boosting Multi-View Indoor 3D Object Detection Via Adaptive 3D Volume ConstructionRunmin Zhang, Zhu Yu, Si-Yuan Cao, Lingyu Zhu et al.ICCV 2025 · 3 citations
- DVGT: Driving Visual Geometry TransformerSicheng Zuo, Zixun Xie, Wenzhao Zheng, Shaoqing Xu et al.CVPR 2026 · 23 citations
- UniDet3D: Multi-dataset Indoor 3D Object DetectionMaksim Kolodiazhnyi, Anna Vorontsova, Matvey Skripkin, Danila Rukhovich et al.AAAI 2025 · 7 citations
