Learning Long-range Information with Dual-Scale Transformers for Indoor Scene Completion
Ziqi Wang, Fei Luo, Xiaoxiao Long, Wenxiao Zhang, Chunxia Xiao
Abstract
Due to the limited resolution of 3D sensors and the inevitable mutual occlusion between objects, 3D scans of real scenes are commonly incomplete. Previous scene completion methods struggle to capture long-range spatial context, resulting in unsatisfactory completion results. To alleviate the problem, we propose a novel Dual-Scale Transformer Network (DST-Net) that efficiently utilizes both long-range and short-range spatial context information to improve the quality of 3D scene completion. To reduce the heavy computation cost of extracting long-range features via transformers, DST-Net adopts a self-supervised two-stage completion strategy. In the first stage, we split the input scene into blocks and perform completion on individual blocks. In the second stage, the blocks are merged together as a whole and then further refined to improve completeness. More importantly, we propose a contrastive attention training strategy to encourage the transformers to learn distinguishable features for better scene completion. Experiments on datasets of Matterport3D, ScanNet, and ICL-NUIM demonstrate that our method can generate better completion results, and our method outperforms the state-of-the-art methods quantitatively and qualitatively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 79bed42a-8442-4a49-bc56-89e63595b6e3Builds on10
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- SnowflakeNet: Point Cloud Completion by Snowflake Point Deconvolution with Skip-TransformerPeng Xiang, Xin Wen, Yu-Shen Liu, Yan-Pei Cao et al.ICCV 2021 · 318 citations
- Neural RGB-D Surface ReconstructionDejan Azinovic, Ricardo Martin-Brualla, Dan B. Goldman, Matthias Nießner et al.CVPR 2022 · 272 citations
Related papers
- Point Cloud Semantic Scene Completion from RGB-D ImagesShoulong Zhang, Shuai Li, Aimin Hao, Hong QinAAAI 2021 · 13 citations
- Monocular Scene Reconstruction with 3D SDF TransformersWeihao Yuan, Xiaodong Gu, Heng Li, Zilong Dong et al.ICLR 2023 · 4 citations
- Learning Local Displacements for Point Cloud CompletionYida Wang, David Joseph Tan, Nassir Navab, Federico TombariCVPR 2022 · 58 citations
- SD-Net: Spatially-Disentangled Point Cloud Completion NetworkJunxian Chen, Ying Liu, Yiqi Liang, Dandan Long et al.ACM MM 2023 · 10 citations
- Deep Point Cloud ReconstructionJaesung Choe, Byeongin Joung, François Rameau, Jaesik Park et al.ICLR 2022 · 28 citations
