Multi-view 3D Reconstruction with Transformers
Dan Wang, Xinrui Cui, Xun Chen, Zhengxia Zou, Tianyang Shi, Septimiu E. Salcudean, Z. Jane Wang, Rabab Ward
Abstract
Deep CNN-based methods have so far achieved the state of the art results in multi-view 3D object reconstruction. Despite the considerable progress, the two core modules of these methods -multi-view feature extraction and fusion, are usually investigated separately, and the object relations in different views are rarely explored. In this paper, inspired by the recent great success in self-attention-based Transformer models, we reformulate the multi-view 3D reconstruction as a sequenceto-sequence prediction problem and propose a new framework named 3D Volume Transformer (VolT) for such a task. Unlike previous CNN-based methods using a separate design, we unify the feature extraction and view fusion in a single Transformer network. A natural advantage of our design lies in the exploration of view-to-view relationships using self-attention among multiple unordered inputs. On ShapeNet -a large-scale 3D reconstruction benchmark dataset, our method achieves a new state-of-the-art accuracy in multi-view reconstruction with fewer parameters (70% less) than other CNN-based methods. Experimental results also suggest the strong scaling capability of our method. Our code will be made publicly available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b4ea0440-cf91-40d9-bbf9-215bd78032f4Cited by top-tier papers20
- SAM 3D: 3Dfy Anything in ImagesXingyu Chen, Fu-Jen Chu, Pierre Gleize, Kevin J Liang et al.CVPR 2026 · 280 citations
- Locally Attentional SDF Diffusion for Controllable 3D Shape GenerationXin-Yang Zheng, Hao Pan, Peng-Shuai Wang, Xin Tong et al.SIGGRAPH 2023 · 122 citations
- Sketch and Text Guided Diffusion Model for Colored Point Cloud GenerationZijie Wu, Yaonan Wang, Mingtao Feng, He Xie et al.ICCV 2023 · 55 citations
- Iterative Superquadric Recomposition of 3D Objects from Multiple ViewsStephan Alaniz, Massimiliano Mancini, Zeynep AkataICCV 2023 · 21 citations
- MV-DeepSDF: Implicit Modeling with Multi-Sweep Point Clouds for 3D Vehicle Reconstruction in Autonomous DrivingYibo Liu, Kelly Zhu, Guile Wu, Yuan Ren et al.ICCV 2023 · 18 citations
Builds on6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu et al.ICML 2020 · 1,773 citations
- Attention is not all you need: pure attention loses rank doubly exponentially with depthYihe Dong, Jean-Baptiste Cordonnier, Andreas LoukasICML 2021 · 522 citations
Related papers
- Pix2Vox: Context-Aware 3D Reconstruction From Single and Multi-View ImagesHaozhe Xie, Hongxun Yao, Xiaoshuai Sun, Shangchen Zhou et al.ICCV 2019 · 373 citations
- UMIFormer: Mining the Correlations between Similar Tokens for Multi-View 3D ReconstructionZhenwei Zhu, Liying Yang, Ning Li, Chaohao Jiang et al.ICCV 2023 · 12 citations
- Long-Range Grouping Transformer for Multi-View 3D ReconstructionLiying Yang, Zhenwei Zhu, Xuxin Lin, Jian Nong et al.ICCV 2023 · 11 citations
- Monocular Scene Reconstruction with 3D SDF TransformersWeihao Yuan, Xiaodong Gu, Heng Li, Zilong Dong et al.ICLR 2023 · 4 citations
- Voxel Set Transformer: A Set-to-Set Approach to 3D Object Detection from Point CloudsChenhang He, Ruihuang Li, Shuai Li, Lei ZhangCVPR 2022 · 217 citations
