Inner-Outer Aware Reconstruction Model for Monocular 3D Scene Reconstruction
Yukun Qiu, Guo-Hao Xu, Wei-Shi Zheng
Abstract
Monocular 3D scene reconstruction aims to reconstruct the 3D structure of scenes based on posed images. Recent volumetric-based methods directly predict the truncated signed distance function (TSDF) volume and have achieved promising results. The memory cost of volumetric-based methods will grow cubically as the volume size increases, so a coarse-to-fine strategy is necessary for saving memory. Specifically, the coarse-to-fine strategy distinguishes surface voxels from non-surface voxels, and only potential surface voxels are considered in the succeeding procedure. However, the non-surface voxels have various features, and in particular, the voxels on the inner side of the surface are quite different from those on the outer side since there exists an intrinsic gap between them. Therefore, grouping inner-surface and outer-surface voxels into the same class will force the classifier to spend its capacity to bridge the gap. By contrast, it is relatively easy for the classifier to distinguish inner-surface and outer-surface voxels due to the intrinsic gap. Inspired by this, we propose the inner-outer aware reconstruction (IOAR) model. IOAR explores a new coarse-to-fine strategy to classify outer-surface, inner-surface and surface voxels. In addition, IOAR separates occupancy branches from TSDF branches to avoid mutual interference between them. Since our model can better classify the surface, outer-surface and inner-surface voxels, it can predict more precise meshes than existing methods. Experiment results on ScanNet, ICL-NUIM and TUM-RGBD datasets demonstrate the effectiveness and generalization of our model. The code is available at https: //github.com/YorkQiu/InnerOuterAwareReconstruction .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on9
- Neural Sparse Voxel FieldsLingjie Liu, Jiatao Gu, Kyaw Zaw Lin, Tat-Seng Chua et al.NeurIPS 2020 · 1,535 citations
- TransformerFusion: Monocular RGB Scene Reconstruction using TransformersAljaz Bozic, Pablo R. Palafox, Justus Thies, Angela Dai et al.NeurIPS 2021 · 185 citations
- Multi-View Stereo by Temporal Nonparametric FusionYuxin Hou, Juho Kannala, Arno SolinICCV 2019 · 99 citations
- NeRFusion: Fusing Radiance Fields for Large-Scale Scene ReconstructionXiaoshuai Zhang, Sai Bi, Kalyan Sunkavalli, Hao Su et al.CVPR 2022 · 94 citations
- Cascade Cost Volume for High-Resolution Multi-View Stereo and Stereo MatchingXiaodong Gu, Zhiwen Fan, Siyu Zhu, Zuozhuo Dai et al.CVPR 2020
Related papers
- FineRecon: Depth-aware Feed-forward Network for Detailed 3D ReconstructionNoah Stier, Anurag Ranjan, Alex Colburn, Yajie Yan et al.ICCV 2023 · 29 citations
- Monocular Scene Reconstruction with 3D SDF TransformersWeihao Yuan, Xiaodong Gu, Heng Li, Zilong Dong et al.ICLR 2023 · 4 citations
- Behind the Veil: Enhanced Indoor 3D Scene Reconstruction with Occluded Surfaces CompletionSu Sun, Cheng Zhao, Yuliang Guo, Ruoyu Wang et al.CVPR 2024 · 3 citations
- NeuralRecon: Real-Time Coherent 3D Reconstruction From Monocular VideoJiaming Sun, Yiming Xie, Linghao Chen, Xiaowei Zhou et al.CVPR 2021
- MonoNeRD: NeRF-like Representations for Monocular 3D Object DetectionJunkai Xu, Liang Peng, Haoran Chen, Hao Li et al.ICCV 2023 · 54 citations
