Lune

NeurIPS2025Top-tier venue

How Many Tokens Do 3D Point Cloud Transformer Architectures Really Need?

Tuan Anh Tran, Duy M. H. Nguyen, Hoai-Chau Tran, Michael Barz, Khoa D. Doan, Roger Wattenhofer, Ngo Anh Vien, Mathias Niepert, Daniel Sonntag, Paul Swoboda

2025Year
5Citations
1Top-tier citations

Abstract

Recent advances in 3D point cloud transformers have led to state-of-the-art results in tasks such as semantic segmentation and reconstruction. However, these models typically rely on dense token representations, incurring high computational and memory costs during training and inference. In this work, we present the finding that tokens are remarkably redundant, leading to substantial inefficiency. We introduce GitMerge3D, a globally informed graph token merging method that can reduce the token count by up to 90-95% while maintaining competitive performance. This finding challenges the prevailing assumption that more tokens inherently yield better performance and highlights that many current models are over-tokenized and under-optimized for scalability. We validate our method across multiple 3D vision tasks and show consistent improvements in computational efficiency. This work is the first to assess redundancy in large-scale 3D transformer models, providing insights into the development of more efficient 3D foundation architectures. Our code and checkpoints are publicly available at https://gitmerge3d.github.io.

• Through a systematic study, we uncover a surprising degree of token redundancy in stateof-the-art point cloud transformers, showing that up to 90-95% of tokens can be removed without significant performance drop. This challenges the common assumption that dense tokenization is essential for 3D transformer effectiveness.

• We propose a 3D-specific token merging strategy that integrates local geometric structure and attention saliency to estimate voxel importance, enabling aggressive token reduction with minimal accuracy degradation.

• We validate the proposed token merging across 3D semantic segmentation, reconstruction, and detection tasks, achieving substantial efficiency gains and, in some cases, surpassing the baseline performance with minimal fine-tuning. We expect these findings will pave the way for future research toward lightweight and scalable transformer architectures for 3D point cloud processing.

2 Related Work 3D Point Cloud Architectures. 3D point cloud understanding has evolved through multiple architectural paradigms. Early approaches include projection-based methods [13, 48, 49, 76], which project point clouds onto 2D image planes for processing with standard CNNs, and voxel-based methods [63, 75, 15, 32, 88], which discretize space into regular grids to apply 3D convolutions. While effective, these techniques often suffer from resolution loss, high memory consumption, or limited geometric expressiveness. To address these challenges, point-based methods such as PointNet [67], PointNet++ [68], and the more recent PointMLP [60] directly operate on raw point sets, preserving fine-grained spatial structure. However, their reliance on local operations can still limit global context modeling. This has led to a growing shift toward transformer-based architectures that better capture long-range dependencies in 3D data. The Point Transformer (PTv) family, spanning PTv-1 [102], PTv-2 [93] to the more scalable PTv3 [92], adapts attention mechanisms to unordered point sets and has become a state-of-the-art backbone for 3D tasks such as semantic segmentation [47, 87, 97, 46], object detection [25, 36, 57, 92], and reconstruction [44, 11, 80, 14]. Building on PTv3's success, recent variants like PTv3 Sonata [91], pretrained on 140K point clouds for improved generalization, and Splatformer [14], tailored for robust 3D novel view synthesis, further demonstrate the versatility and dominance of transformer-based models in modern 3D vision pipelines.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 39858205-0da5-4abe-bf92-3b1e67e240fb

Cited by top-tier papers1

Ask how each one uses it

Builds on51

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines