ImGeoNet: Image-induced Geometry-aware Voxel Representation for Multi-view 3D Object Detection
Tao Tu, Shun-Po Chuang, Yu-Lun Liu, Cheng Sun, Ke Zhang, Donna Roy, Cheng-Hao Kuo, Min Sun
Abstract
We propose ImGeoNet, a multi-view image-based 3D object detection framework that models a 3D space by an image-induced geometry-aware voxel representation. Unlike previous methods which aggregate 2D features into 3D voxels without considering geometry, ImGeoNet learns to induce geometry from multi-view images to alleviate the confusion arising from voxels of free space, and during the inference phase, only images from multiple views are required. Besides, a powerful pre-trained 2D feature extractor can be leveraged by our representation, leading to a more robust performance. To evaluate the effectiveness of ImGeoNet, we conduct quantitative and qualitative experiments on three indoor datasets, namely ARKitScenes, ScanNetV2, and ScanNet200. The results demonstrate that ImGeoNet outperforms the current state-of-the-art multiview image-based method, ImVoxelNet, on all three datasets in terms of detection accuracy. In addition, ImGeoNet shows great data efficiency by achieving results comparable to ImVoxelNet with 100 views while utilizing only 40 views. Furthermore, our studies indicate that our proposed image-induced geometry-aware representation can enable image-based methods to attain superior detection accuracy than the seminal point cloud-based method, VoteNet, in two practical scenarios: (1) scenarios where point clouds are sparse and noisy, such as in ARKitScenes, and (2) scenarios involve diverse object classes, particularly classes of small objects, as in the case in ScanNet200. Project page: https://ttaoretw.github.io/imgeonet .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- MVSDet: Multi-View Indoor 3D Object Detection via Efficient Plane SweepsYating Xu, Chen Li, Gim Hee LeeNeurIPS 2024 · 10 citations
- Zoo3D: Zero-Shot 3D Object Detection at Scene LevelAndrey Lemeshko, Bulat Gabdullin, Nikita Drozdov, Anton Konushin et al.CVPR 2026 · 5 citations
- Unsupervised Multi-view Pedestrian DetectionMengyin Liu, Chao Zhu, Shiqi Ren, Xu-Cheng YinACM MM 2024 · 4 citations
- CN-RMA: Combined Network with Ray Marching Aggregation for 3D Indoor Object Detection from Multi-View ImagesGuanlin Shen, Jingwei Huang, Zhihua Hu, Bin WangCVPR 2024 · 3 citations
- Voxify3D: Pixel Art Meets Volumetric RenderingYi-Chuan Huang, Jiewen Chan, Hao-Jen Chien, Yu-Lun LiuCVPR 2026 · 3 citations
Builds on34
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 1,467 citations
- Voxel R-CNN: Towards High Performance Voxel-based 3D Object DetectionJiajun Deng, Shaoshuai Shi, Peiwei Li, Wengang Zhou et al.AAAI 2021 · 1,128 citations
- MVSNeRF: Fast Generalizable Radiance Field Reconstruction from Multi-View StereoAnpei Chen, Zexiang Xu, Fuqiang Zhao, Xiaoshuai Zhang et al.ICCV 2021 · 1,024 citations
- Direct Voxel Grid Optimization: Super-fast Convergence for Radiance Fields ReconstructionCheng Sun, Min Sun, Hwann-Tzong ChenCVPR 2022 · 859 citations
Related papers
- OpenM3D: Open Vocabulary Multi-View Indoor 3D Object Detection without Human AnnotationsPeng-Hao Hsu, Ke Zhang, Fu-En Wang, Tao Tu et al.ICCV 2025 · 3 citations
- Boosting Multi-View Indoor 3D Object Detection Via Adaptive 3D Volume ConstructionRunmin Zhang, Zhu Yu, Si-Yuan Cao, Lingyu Zhu et al.ICCV 2025 · 3 citations
- ImVoteNet: Boosting 3D Object Detection in Point Clouds With Image VotesCharles R. Qi, Xinlei Chen, Or Litany, Leonidas J. GuibasCVPR 2020
- VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object DetectionYang Cao, Feize Wu, Dave Chen, Yingji Zhong et al.CVPR 2026 · 6 citations
- NeRF-Det: Learning Geometry-Aware Volumetric Representation for Multi-View 3D Object DetectionChenfeng Xu, Bichen Wu, Ji Hou, Sam S. Tsai et al.ICCV 2023 · 71 citations
