AutoRecon: Automated 3D Object Discovery and Reconstruction
Yuang Wang, Xingyi He, Sida Peng, Haotong Lin, Hujun Bao, Xiaowei Zhou
Abstract
A fully automated object reconstruction pipeline is crucial for digital content creation. While the area of 3D reconstruction has witnessed profound developments, the removal of background to obtain a clean object model still relies on different forms of manual labor, such as bounding box labeling, mask annotations, and mesh manipulations. In this paper, we propose a novel framework named AutoRecon for the automated discovery and reconstruction of an object from multi-view images. We demonstrate that foreground objects can be robustly located and segmented from SfM point clouds by leveraging self-supervised 2D vision transformer features. Then, we reconstruct decomposed neural scene representations with dense supervision provided by the decomposed point clouds, resulting in accurate object reconstruction and segmentation. Experiments on the DTU, BlendedMVS and CO3D-V2 datasets demonstrate the effectiveness and robustness of AutoRecon. The code and supplementary material are available on the project page: https://zju3dv.github.io/autorecon/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Language Embedded 3D Gaussians for Open-Vocabulary Scene UnderstandingJin-Chuan Shi, Miao Wang, Hao-Bin Duan, Shao-Hua GuanCVPR 2024 · 57 citations
- ObjectSDF++: Improved Object-Compositional Neural Implicit SurfacesQianyi Wu, Kaisiyuan Wang, Kejie Li, Jianmin Zheng et al.ICCV 2023 · 44 citations
- NeuSurf: On-Surface Priors for Neural Surface Reconstruction from Sparse Input ViewsHan Huang, Yulun Wu, Junsheng Zhou, Ge Gao et al.AAAI 2024 · 42 citations
- VoMP: Predicting Volumetric Mechanical Property FieldsRishit Dagli, Donglai Xiang, Vismay Modi, Charles Loop et al.ICLR 2026 · 13 citations
- Multi-View Aggregation Network for Dichotomous Image SegmentationQian Yu, Xiaoqi Zhao, Youwei Pang, Lihe Zhang et al.CVPR 2024 · 13 citations
Builds on30
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
- NeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view ReconstructionPeng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt et al.NeurIPS 2021 · 2,500 citations
- Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance FieldsJonathan T. Barron, Ben Mildenhall, Dor Verbin, Pratul P. Srinivasan et al.CVPR 2022 · 1,603 citations
Related papers
- Multi-view 3D Reconstruction with TransformersDan Wang, Xinrui Cui, Xun Chen, Zhengxia Zou et al.ICCV 2021 · 111 citations
- Reliable-View 2D-3D Key-Part Aligned Transformer with Reinforced Masking for 3D Point Cloud UnderstandingXianglong Jin, Zheng Wang, Rong Wang, Feiping NieAAAI 2026
- Rayzer: a Self-Supervised Large View Synthesis ModelHanwen Jiang, Hao Tan, Peng Wang, Hai Jin et al.ICCV 2025 · 12 citations
- Discovering 3D Parts from Image CollectionsChun-Han Yao, Wei-Chih Hung, Varun Jampani, Ming-Hsuan YangICCV 2021 · 21 citations
- FoundObj: Self-supervised Foundation Models as Rewards for Label-free 3D Object SegmentationZihui Zhang, Zhixuan Sun, Yafei YANG, Jinxi Li et al.ICML 2026
