PatchFormer: An Efficient Point Transformer with Patch Attention
Cheng Zhang, Haocheng Wan, Xinyi Shen, Zizhao Wu
Abstract
The point cloud learning community witnesses a modeling shift from CNNs to Transformers, where pure Transformer architectures have achieved top accuracy on the major learning benchmarks. However, existing point Transformers are computationally expensive since they need to generate a large attention map, which has quadratic complexity (both in space and time) with respect to input size. To solve this shortcoming, we introduce Patch ATtention (PAT) to adaptively learn a much smaller set of bases upon which the attention maps are computed. By a weighted summation upon these bases, PAT not only captures the global shape context but also achieves linear complexity to input size. In addition, we propose a lightweight Multi-Scale aTtention (MST) block to build attentions among features of different scales, providing the model with multi-scale features. Equipped with the PAT and MST, we construct our neural architecture called PatchFormer that integrates both modules into a joint framework for point cloud learning. Extensive experiments demonstrate that our network achieves comparable accuracy on general point cloud learning tasks with <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> speed-up than previous point Transformers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers21
- E2VPT: An Effective and Efficient Approach for Visual Prompt TuningCheng Han, Qifan Wang, Yiming Cui, Zhiwen Cao et al.ICCV 2023 · 108 citations
- LitePT: Lighter Yet Stronger Point TransformerYuanwen Yue, Damien Robert, Jianyuan Wang, Sunghwan Hong et al.CVPR 2026 · 25 citations
- Improved MLP Point Cloud Processing with High-Dimensional Positional EncodingYanmei Zou, Hongshan Yu, Zhengeng Yang, Zechuan Li et al.AAAI 2024 · 15 citations
- Pamba: Enhancing Global Interaction in Point Clouds via State Space ModelZhuoyuan Li, Yubo Ai, Jiahao Lu, Chuxin Wang et al.AAAI 2025 · 12 citations
- RadarMOSEVE: A Spatial-Temporal Transformer Network for Radar-Only Moving Object Segmentation and Ego-Velocity EstimationChangsong Pang, Xieyuanli Chen, Yimin Liu, Huimin Lu et al.AAAI 2024 · 8 citations
Builds on18
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- KPConv: Flexible and Deformable Convolution for Point CloudsHugues Thomas, Charles R. Qi, Jean-Emmanuel Deschaud, Beatriz Marcotegui et al.ICCV 2019 · 3,193 citations
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 2,072 citations
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei et al.CVPR 2022 · 1,847 citations
Related papers
- OctFormer: Octree-based Transformers for 3D Point CloudsPeng-Shuai WangSIGGRAPH 2023 · 123 citations
- 3D Object Detection With PointformerXuran Pan, Zhuofan Xia, Shiji Song, Li Erran Li et al.CVPR 2021
- ConDaFormer: Disassembled Transformer with Local Structure Enhancement for 3D Point Cloud UnderstandingLunhao Duan, Shanshan Zhao, Nan Xue, Mingming Gong et al.NeurIPS 2023 · 37 citations
- Self-Positioning Point-Based Transformer for Point Cloud UnderstandingJinyoung Park, Sanghyeok Lee, Sihyeon Kim, Yunyang Xiong et al.CVPR 2023
- PointConvFormer: Revenge of the Point-based ConvolutionWenxuan Wu, Fuxin Li, Qi ShanCVPR 2023
