How Many Tokens Do 3D Point Cloud Transformer Architectures Really Need?
Tuan Anh Tran, Duy M. H. Nguyen, Hoai-Chau Tran, Michael Barz, Khoa D. Doan, Roger Wattenhofer, Ngo Anh Vien, Mathias Niepert, Daniel Sonntag, Paul Swoboda
Abstract
Recent advances in 3D point cloud transformers have led to state-of-the-art results in tasks such as semantic segmentation and reconstruction. However, these models typically rely on dense token representations, incurring high computational and memory costs during training and inference. In this work, we present the finding that tokens are remarkably redundant, leading to substantial inefficiency. We introduce GitMerge3D, a globally informed graph token merging method that can reduce the token count by up to 90-95% while maintaining competitive performance. This finding challenges the prevailing assumption that more tokens inherently yield better performance and highlights that many current models are over-tokenized and under-optimized for scalability. We validate our method across multiple 3D vision tasks and show consistent improvements in computational efficiency. This work is the first to assess redundancy in large-scale 3D transformer models, providing insights into the development of more efficient 3D foundation architectures. Our code and checkpoints are publicly available at https://gitmerge3d.github.io.
• Through a systematic study, we uncover a surprising degree of token redundancy in stateof-the-art point cloud transformers, showing that up to 90-95% of tokens can be removed without significant performance drop. This challenges the common assumption that dense tokenization is essential for 3D transformer effectiveness.
• We propose a 3D-specific token merging strategy that integrates local geometric structure and attention saliency to estimate voxel importance, enabling aggressive token reduction with minimal accuracy degradation.
• We validate the proposed token merging across 3D semantic segmentation, reconstruction, and detection tasks, achieving substantial efficiency gains and, in some cases, surpassing the baseline performance with minimal fine-tuning. We expect these findings will pave the way for future research toward lightweight and scalable transformer architectures for 3D point cloud processing.
2 Related Work 3D Point Cloud Architectures. 3D point cloud understanding has evolved through multiple architectural paradigms. Early approaches include projection-based methods [13, 48, 49, 76], which project point clouds onto 2D image planes for processing with standard CNNs, and voxel-based methods [63, 75, 15, 32, 88], which discretize space into regular grids to apply 3D convolutions. While effective, these techniques often suffer from resolution loss, high memory consumption, or limited geometric expressiveness. To address these challenges, point-based methods such as PointNet [67], PointNet++ [68], and the more recent PointMLP [60] directly operate on raw point sets, preserving fine-grained spatial structure. However, their reliance on local operations can still limit global context modeling. This has led to a growing shift toward transformer-based architectures that better capture long-range dependencies in 3D data. The Point Transformer (PTv) family, spanning PTv-1 [102], PTv-2 [93] to the more scalable PTv3 [92], adapts attention mechanisms to unordered point sets and has become a state-of-the-art backbone for 3D tasks such as semantic segmentation [47, 87, 97, 46], object detection [25, 36, 57, 92], and reconstruction [44, 11, 80, 14]. Building on PTv3's success, recent variants like PTv3 Sonata [91], pretrained on 140K point clouds for improved generalization, and Splatformer [14], tailored for robust 3D novel view synthesis, further demonstrate the versatility and dominance of transformer-based models in modern 3D vision pipelines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 39858205-0da5-4abe-bf92-3b1e67e240fbCited by top-tier papers1
Ask how each one uses itBuilds on51
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei et al.CVPR 2022 · 1,847 citations
Related papers
- Point Transformer V3: Simpler, Faster, StrongerXiaoyang Wu, Li Jiang, Peng-Shuai Wang, Zhijian Liu et al.CVPR 2024
- LiteVGGT: Boosting Vanilla VGGT via Geometry-aware Cached Token MergingZhijian Shu, Cheng Lin, Tao Xie, Wei Yin et al.CVPR 2026 · 17 citations
- OctFormer: Octree-based Transformers for 3D Point CloudsPeng-Shuai WangSIGGRAPH 2023 · 123 citations
- FastVGGT: Fast Visual Geometry TransformerYou Shen, Zhipeng Zhang, Yansong Qu, Xiawu Zheng et al.ICLR 2026 · 73 citations
- Efficient 3D Semantic Segmentation with Superpoint TransformerDamien Robert, Hugo Raguet, Loïc LandrieuICCV 2023 · 131 citations
