Point Transformer V2: Grouped Vector Attention and Partition-based Pooling
Xiaoyang Wu, Yixing Lao, Li Jiang, Xihui Liu, Hengshuang Zhao
Abstract
As a pioneering work exploring transformer architecture for 3D point cloud understanding, Point Transformer achieves impressive results on multiple highly competitive benchmarks. In this work, we analyze the limitations of the Point Transformer and propose our powerful and efficient Point Transformer V2 model with novel designs that overcome the limitations of previous work. In particular, we first propose group vector attention, which is more effective than the previous version of vector attention. Inheriting the advantages of both learnable weight encoding and multi-head attention, we present a highly effective implementation of grouped vector attention with a novel grouped weight encoding layer. We also strengthen the position information for attention by an additional position encoding multiplier. Furthermore, we design novel and lightweight partition-based pooling methods which enable better spatial alignment and more efficient sampling. Extensive experiments show that our model achieves better performance than its predecessor and achieves state-of-the-art on several challenging 3D point cloud understanding benchmarks, including 3D point cloud segmentation on ScanNet v2 and S3DIS and 3D point cloud classification on ModelNet40. Our code will be available at https://github.com/Gofinge/PointTransformerV2 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers176
- PointMamba: A Simple State Space Model for Point Cloud AnalysisDingkang Liang, Xin Zhou, Wei Xu, Xingkui Zhu et al.NeurIPS 2024 · 380 citations
- OctFormer: Octree-based Transformers for 3D Point CloudsPeng-Shuai WangSIGGRAPH 2023 · 123 citations
- Point Cloud Mamba: Point Cloud Learning via State Space ModelTao Zhang, Haobo Yuan, Lu Qi, Jiangning Zhang et al.AAAI 2025 · 110 citations
- SpatialLM: Training Large Language Models for Structured Indoor ModelingYongsen Mao, Junhao Zhong, Chuan Fang, Jia Zheng et al.NeurIPS 2025 · 89 citations
- Mamba3D: Enhancing Local Features for 3D Point Cloud Analysis via State Space ModelXu Han, Yuan Tang, Zhaoxuan Wang, Xianzhi LiACM MM 2024 · 86 citations
Builds on17
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- KPConv: Flexible and Deformable Convolution for Point CloudsHugues Thomas, Charles R. Qi, Jean-Emmanuel Deschaud, Beatriz Marcotegui et al.ICCV 2019 · 3,193 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
Related papers
- Stratified Transformer for 3D Point Cloud SegmentationXin Lai, Jianhui Liu, Li Jiang, Liwei Wang et al.CVPR 2022 · 494 citations
- Few-Shot 3D Point Cloud Semantic Segmentation via Stratified Class-Specific Attention Based Transformer NetworkCanyu Zhang, Zhenyao Wu, Xinyi Wu, Ziyu Zhao et al.AAAI 2023 · 31 citations
- Point TransformerHengshuang Zhao, Li Jiang, Jiaya Jia, Philip H. S. Torr et al.ICCV 2021 · 23 citations
- Fast Point TransformerChunghyun Park, Yoonwoo Jeong, Minsu Cho, Jaesik ParkCVPR 2022
- Efficient 3D Semantic Segmentation with Superpoint TransformerDamien Robert, Hugo Raguet, Loïc LandrieuICCV 2023 · 131 citations
