ConDaFormer: Disassembled Transformer with Local Structure Enhancement for 3D Point Cloud Understanding
Lunhao Duan, Shanshan Zhao, Nan Xue, Mingming Gong, Gui-Song Xia, Dacheng Tao
摘要
Transformers have been recently explored for 3D point cloud understanding with impressive progress achieved. A large number of points, over 0.1 million, make the global self-attention infeasible for point cloud data. Thus, most methods propose to apply the transformer in a local region, e.g., spherical or cubic window. However, it still contains a large number of Query-Key pairs, which require high computational costs. In addition, previous methods usually learn the query, key, and value using a linear projection without modeling the local 3D geometric structure. In this paper, we attempt to reduce the costs and model the local geometry prior by developing a new transformer block, named ConDaFormer ‡ . Technically, ConDaFormer disassembles the cubic window into three orthogonal 2D planes, leading to fewer points when modeling the attention in a similar range. The disassembling operation is beneficial to enlarging the range of attention without increasing the computational complexity but ignores some contexts. To provide a remedy, we develop a local structure enhancement strategy that introduces a depth-wise convolution before and after the attention. This scheme can also capture the local geometric information. Taking advantage of these designs, ConDaFormer captures both long-range contextual information and local priors. The effectiveness is demonstrated by experimental results on several 3D point cloud understanding benchmarks. Our code will be available at https://github.com/LHDuan/ConDaFormer .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- LitePT: Lighter Yet Stronger Point TransformerYuanwen Yue, Damien Robert, Jianyuan Wang, Sunghwan Hong 等CVPR 2026 · 被引用 25 次
- UniMix: Towards Domain Adaptive and Generalizable LiDAR Semantic Segmentation in Adverse WeatherHaimei Zhao, Jing Zhang, Zhuo Chen, Shanshan Zhao 等CVPR 2024 · 被引用 25 次
- Pamba: Enhancing Global Interaction in Point Clouds via State Space ModelZhuoyuan Li, Yubo Ai, Jiahao Lu, Chuxin Wang 等AAAI 2025 · 被引用 12 次
- LinNet: Linear Network for Efficient Point Cloud Representation LearningHao Deng, Kunlei Jing, Shengmei Chen, Cheng Liu 等NeurIPS 2024 · 被引用 12 次
- On-the-fly Point Feature Representation for Point Clouds AnalysisJiangyi Wang, Zhongyao Cheng, Na Zhao, Jun Cheng 等ACM MM 2024 · 被引用 9 次
它引用的顶会 Paper48
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- KPConv: Flexible and Deformable Convolution for Point CloudsHugues Thomas, Charles R. Qi, Jean-Emmanuel Deschaud, Beatriz Marcotegui 等ICCV 2019 · 被引用 3,193 次
- CCNet: Criss-Cross Attention for Semantic SegmentationZilong Huang, Xinggang Wang, Lichao Huang, Chang Huang 等ICCV 2019 · 被引用 2,972 次
- CvT: Introducing Convolutions to Vision TransformersHaiping Wu, Bin Xiao, Noel Codella, Mengchen Liu 等ICCV 2021 · 被引用 2,397 次
相关 Paper
- OctFormer: Octree-based Transformers for 3D Point CloudsPeng-Shuai WangSIGGRAPH 2023 · 被引用 123 次
- PointCFormer: A Relation-Based Progressive Feature Extraction Network for Point Cloud CompletionYi Zhong, Weize Quan, Dong-Ming Yan, Jie Jiang 等AAAI 2025 · 被引用 3 次
- PatchFormer: An Efficient Point Transformer with Patch AttentionCheng Zhang, Haocheng Wan, Xinyi Shen, Zizhao WuCVPR 2022 · 被引用 77 次
- FlatFormer: Flattened Window Attention for Efficient Point Cloud TransformerZhijian Liu, Xinyu Yang, Haotian Tang, Shang Yang 等CVPR 2023
- 3D Object Detection With PointformerXuran Pan, Zhuofan Xia, Shiji Song, Li Erran Li 等CVPR 2021
