Multi-view 3D Reconstruction with Transformers
Dan Wang, Xinrui Cui, Xun Chen, Zhengxia Zou, Tianyang Shi, Septimiu E. Salcudean, Z. Jane Wang, Rabab Ward
摘要
Deep CNN-based methods have so far achieved the state of the art results in multi-view 3D object reconstruction. Despite the considerable progress, the two core modules of these methods -multi-view feature extraction and fusion, are usually investigated separately, and the object relations in different views are rarely explored. In this paper, inspired by the recent great success in self-attention-based Transformer models, we reformulate the multi-view 3D reconstruction as a sequenceto-sequence prediction problem and propose a new framework named 3D Volume Transformer (VolT) for such a task. Unlike previous CNN-based methods using a separate design, we unify the feature extraction and view fusion in a single Transformer network. A natural advantage of our design lies in the exploration of view-to-view relationships using self-attention among multiple unordered inputs. On ShapeNet -a large-scale 3D reconstruction benchmark dataset, our method achieves a new state-of-the-art accuracy in multi-view reconstruction with fewer parameters (70% less) than other CNN-based methods. Experimental results also suggest the strong scaling capability of our method. Our code will be made publicly available.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- SAM 3D: 3Dfy Anything in ImagesXingyu Chen, Fu-Jen Chu, Pierre Gleize, Kevin J Liang 等CVPR 2026 · 被引用 280 次
- Locally Attentional SDF Diffusion for Controllable 3D Shape GenerationXin-Yang Zheng, Hao Pan, Peng-Shuai Wang, Xin Tong 等SIGGRAPH 2023 · 被引用 122 次
- Sketch and Text Guided Diffusion Model for Colored Point Cloud GenerationZijie Wu, Yaonan Wang, Mingtao Feng, He Xie 等ICCV 2023 · 被引用 55 次
- Iterative Superquadric Recomposition of 3D Objects from Multiple ViewsStephan Alaniz, Massimiliano Mancini, Zeynep AkataICCV 2023 · 被引用 21 次
- MV-DeepSDF: Implicit Modeling with Multi-Sweep Point Clouds for 3D Vehicle Reconstruction in Autonomous DrivingYibo Liu, Kelly Zhu, Guile Wu, Yuan Ren 等ICCV 2023 · 被引用 18 次
它引用的顶会 Paper6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu 等ICML 2020 · 被引用 1,773 次
- Attention is not all you need: pure attention loses rank doubly exponentially with depthYihe Dong, Jean-Baptiste Cordonnier, Andreas LoukasICML 2021 · 被引用 522 次
相关 Paper
- Pix2Vox: Context-Aware 3D Reconstruction From Single and Multi-View ImagesHaozhe Xie, Hongxun Yao, Xiaoshuai Sun, Shangchen Zhou 等ICCV 2019 · 被引用 373 次
- UMIFormer: Mining the Correlations between Similar Tokens for Multi-View 3D ReconstructionZhenwei Zhu, Liying Yang, Ning Li, Chaohao Jiang 等ICCV 2023 · 被引用 12 次
- Long-Range Grouping Transformer for Multi-View 3D ReconstructionLiying Yang, Zhenwei Zhu, Xuxin Lin, Jian Nong 等ICCV 2023 · 被引用 11 次
- Monocular Scene Reconstruction with 3D SDF TransformersWeihao Yuan, Xiaodong Gu, Heng Li, Zilong Dong 等ICLR 2023 · 被引用 4 次
- Voxel Set Transformer: A Set-to-Set Approach to 3D Object Detection from Point CloudsChenhang He, Ruihuang Li, Shuai Li, Lei ZhangCVPR 2022 · 被引用 217 次
