Long-Range Grouping Transformer for Multi-View 3D Reconstruction
Liying Yang, Zhenwei Zhu, Xuxin Lin, Jian Nong, Yanyan Liang
摘要
Nowadays, transformer networks have demonstrated superior performance in many computer vision tasks. In a multi-view 3D reconstruction algorithm following this paradigm, self-attention processing has to deal with intricate image tokens including massive information when facing heavy amounts of view input. The curse of information content leads to the extreme difficulty of model learning. To alleviate this problem, recent methods compress the token number representing each view or discard the attention operations between the tokens from different views. Obviously, they give a negative impact on performance. Therefore, we propose long-range grouping attention (LGA) based on the divide-and-conquer principle. Tokens from all views are grouped for separate attention operations. The tokens in each group are sampled from all views and can provide macro representation for the resided view. The richness of feature learning is guaranteed by the diversity among different groups. An effective and efficient encoder can be established which connects inter-view features using LGA and extract intra-view features using the standard self-attention layer. Moreover, a novel progressive upsampling decoder is also designed for voxel generation with relatively high resolution. Hinging on the above, we construct a powerful transformer-based network, called LRGT. Experimental results on ShapeNet verify our method achieves SOTA accuracy in multi-view reconstruction. Code will be available at https://github.com/LiyingCV/ Long-Range-Grouping-Transformer .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- View Transformation Robustness for Multi-View 3D Object Reconstruction with Reconstruction Error-Guided View SelectionQi Zhang, Zhouhang Luo, Tao Yu, Hui HuangAAAI 2025 · 被引用 1 次
- Not All Frame Features are Equal: Video-to-4D Generation via Decoupling Dynamic-Static FeaturesLiying Yang, Chen Liu, Zhenwei Zhu, Ajian Liu 等ICCV 2025
- Structured Diffusion Bridges: Inductive Bias for Denoising Diffusion BridgesEitan Kosman, Gabriele Serussi, Chaim BaskinICML 2026
- PointCHR: Point Cloud Analysis via Curvature-Aware Hyperbolic RectificationXinxing Yu, Liying Yang, Hao Mo, Hui Ma 等ICML 2026
- Cross-Modal 3D Representation with Multi-View Images and Point CloudsZiyang Zhou, Pinghui Wang, Zi Liang, Haitao Bai 等CVPR 2025
它引用的顶会 Paper9
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped WindowsXiaoyi Dong, Jianmin Bao, Dongdong Chen, Weiming Zhang 等CVPR 2022 · 被引用 1,207 次
- How Do Vision Transformers Work?Namuk Park, Songkuk KimICLR 2022 · 被引用 653 次
- Pix2Vox: Context-Aware 3D Reconstruction From Single and Multi-View ImagesHaozhe Xie, Hongxun Yao, Xiaoshuai Sun, Shangchen Zhou 等ICCV 2019 · 被引用 373 次
相关 Paper
- Multi-view 3D Reconstruction with TransformersDan Wang, Xinrui Cui, Xun Chen, Zhengxia Zou 等ICCV 2021 · 被引用 111 次
- UMIFormer: Mining the Correlations between Similar Tokens for Multi-View 3D ReconstructionZhenwei Zhu, Liying Yang, Ning Li, Chaohao Jiang 等ICCV 2023 · 被引用 12 次
- CATANet: Efficient Content-Aware Token Aggregation for Lightweight Image Super-ResolutionXin Liu, Jie Liu, Jie Tang, Gangshan WuCVPR 2025
- FlashVGGT: Efficient and Scalable Visual Geometry Transformers with Compressed Descriptor AttentionZipeng Wang, Dan XuCVPR 2026 · 被引用 14 次
- Learning Compact 3D Representations from Feed-Forward Novel View SynthesisHonggyu An, Jaewoo Jung, Mungyeom Kim, Chaehyun Kim 等CVPR 2026
