Geodesic Self-Attention for 3D Point Clouds
Zhengyu Li, Xuan Tang, Zihao Xu, Xihao Wang, Hui Yu, Mingsong Chen, Xian Wei
Abstract
Due to the outstanding competence in capturing long-range relationships, self-attention mechanism has achieved remarkable progress in point cloud tasks. Never-theless, point cloud object often has complex non-Euclidean spatial structures, with the behavior changing dynamically and unpredictably. Most current self-attention modules highly rely on the dot product multiplication in Euclidean space, which cannot capture internal non-Euclidean structures of point cloud objects, especially the long-range relationships along the curve of the implicit manifold surface represented by point cloud objects. To address this problem, in this paper, we introduce a novel metric on the Riemannian manifold to capture the long-range geometrical dependencies of point cloud objects to replace traditional self-attention modules, namely, the G eodesic S elf-A ttention (GSA) module. Our approach achieves state-of-the-art performance compared to point cloud Transformers [13, 10, 44, 26] on object classification, few-shot classification and part segmentation benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a47b8065-4447-41b2-aead-7e893aa35976Cited by top-tier papers3
- NeuroGF: A Neural Representation for Fast Geodesic Distance and Path QueriesQijian Zhang, Junhui Hou, Yohanes Yudhi Adikusuma, Wenping Wang et al.NeurIPS 2023 · 11 citations
- GeoMM: On Geodesic Perspective for Multi-modal LearningShibin Mei, Hang Wang, Bingbing NiCVPR 2025
- Curvature-Aware Captioning: Leveraging Geodesic Attention for 3D Scene UnderstandingZiyao He, Yingjie Liu, Yangrui Zhang, Mingsong Chen et al.CVPR 2026
Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- DeepGCNs: Can GCNs Go As Deep As CNNs?Guohao Li, Matthias Müller, Ali K. Thabet, Bernard GhanemICCV 2019 · 1,586 citations
- Voxel R-CNN: Towards High Performance Voxel-based 3D Object DetectionJiajun Deng, Shaoshuai Shi, Peiwei Li, Wengang Zhou et al.AAAI 2021 · 1,128 citations
Related papers
- PointAttN: You Only Need Attention for Point Cloud CompletionJun Wang, Ying Cui, Dongyan Guo, Junxia Li et al.AAAI 2024 · 113 citations
- Point TransformerHengshuang Zhao, Li Jiang, Jiaya Jia, Philip H. S. Torr et al.ICCV 2021 · 23 citations
- Self-Positioning Point-Based Transformer for Point Cloud UnderstandingJinyoung Park, Sanghyeok Lee, Sihyeon Kim, Yunyang Xiong et al.CVPR 2023
- Improving Graph Representation for Point Cloud Segmentation via Attentive FilteringNan Zhang, Zhiyi Pan, Thomas H. Li, Wei Gao et al.CVPR 2023
- DHGCN: Dynamic Hop Graph Convolution Network for Self-Supervised Point Cloud LearningJincen Jiang, Lizhi Zhao, Xuequan Lu, Wei Hu et al.AAAI 2024 · 21 citations
