UMIFormer: Mining the Correlations between Similar Tokens for Multi-View 3D Reconstruction
Zhenwei Zhu, Liying Yang, Ning Li, Chaohao Jiang, Yanyan Liang
Abstract
In recent years, many video tasks have achieved breakthroughs by utilizing the vision transformer and establishing spatial-temporal decoupling for feature extraction. Although multi-view 3D reconstruction also faces multiple images as input, it cannot immediately inherit their success due to completely ambiguous associations between unstructured views. There is not usable prior relationship, which is similar to the temporally-coherence property in a video. To solve this problem, we propose a novel transformer network for Unstructured Multiple Images (UMIFormer). It exploits transformer blocks for decoupled intra-view encoding and designed blocks for token rectification that mine the correlation between similar tokens from different views to achieve decoupled interview encoding. Afterward, all tokens acquired from various branches are compressed into a fixed-size compact representation while preserving rich information for reconstruction by leveraging the similarities between tokens. We empirically demonstrate on ShapeNet and confirm that our decoupled learning method is adaptable for unstructured multiple images. Meanwhile, the experiments also verify our model outperforms existing SOTA methods by a large margin. Code will be available at https://github.com/ GaryZhu1996/UMIFormer .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 77cd83bd-0fed-498a-bb45-1120789c8202Cited by top-tier papers3
- Long-Range Grouping Transformer for Multi-View 3D ReconstructionLiying Yang, Zhenwei Zhu, Xuxin Lin, Jian Nong et al.ICCV 2023 · 11 citations
- View Transformation Robustness for Multi-View 3D Object Reconstruction with Reconstruction Error-Guided View SelectionQi Zhang, Zhouhang Luo, Tao Yu, Hui HuangAAAI 2025 · 1 citation
- Not All Frame Features are Equal: Video-to-4D Generation via Decoupling Dynamic-Static FeaturesLiying Yang, Chen Liu, Zhenwei Zhu, Ajian Liu et al.ICCV 2025
Builds on10
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- Pix2Vox: Context-Aware 3D Reconstruction From Single and Multi-View ImagesHaozhe Xie, Hongxun Yao, Xiaoshuai Sun, Shangchen Zhou et al.ICCV 2019 · 373 citations
Related papers
- Multi-view 3D Reconstruction with TransformersDan Wang, Xinrui Cui, Xun Chen, Zhengxia Zou et al.ICCV 2021 · 111 citations
- SVTformer: Spatial-View-Temporal Transformer for Multi-View 3D Human Pose EstimationWanruo Zhang, Mengyuan Liu, Hong Liu, Wenhao LiAAAI 2025 · 4 citations
- SceneTok: A Compressed, Diffusable Token Space for 3D ScenesMohammad Asim, Christopher Wewer, Jan LenssenCVPR 2026 · 6 citations
- OmniTokenizer: A Joint Image-Video Tokenizer for Visual GenerationJunke Wang, Yi Jiang, Zehuan Yuan, Bingyue Peng et al.NeurIPS 2024 · 132 citations
- UniFormer: Unified Transformer for Efficient Spatial-Temporal Representation LearningKunchang Li, Yali Wang, Peng Gao, Guanglu Song et al.ICLR 2022
