ConDaFormer: Disassembled Transformer with Local Structure Enhancement for 3D Point Cloud Understanding
Lunhao Duan, Shanshan Zhao, Nan Xue, Mingming Gong, Gui-Song Xia, Dacheng Tao
Abstract
Transformers have been recently explored for 3D point cloud understanding with impressive progress achieved. A large number of points, over 0.1 million, make the global self-attention infeasible for point cloud data. Thus, most methods propose to apply the transformer in a local region, e.g., spherical or cubic window. However, it still contains a large number of Query-Key pairs, which require high computational costs. In addition, previous methods usually learn the query, key, and value using a linear projection without modeling the local 3D geometric structure. In this paper, we attempt to reduce the costs and model the local geometry prior by developing a new transformer block, named ConDaFormer ‡ . Technically, ConDaFormer disassembles the cubic window into three orthogonal 2D planes, leading to fewer points when modeling the attention in a similar range. The disassembling operation is beneficial to enlarging the range of attention without increasing the computational complexity but ignores some contexts. To provide a remedy, we develop a local structure enhancement strategy that introduces a depth-wise convolution before and after the attention. This scheme can also capture the local geometric information. Taking advantage of these designs, ConDaFormer captures both long-range contextual information and local priors. The effectiveness is demonstrated by experimental results on several 3D point cloud understanding benchmarks. Our code will be available at https://github.com/LHDuan/ConDaFormer .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5c0b2fcf-a9bd-4080-92e7-7282680e7f36Cited by top-tier papers15
- LitePT: Lighter Yet Stronger Point TransformerYuanwen Yue, Damien Robert, Jianyuan Wang, Sunghwan Hong et al.CVPR 2026 · 25 citations
- UniMix: Towards Domain Adaptive and Generalizable LiDAR Semantic Segmentation in Adverse WeatherHaimei Zhao, Jing Zhang, Zhuo Chen, Shanshan Zhao et al.CVPR 2024 · 25 citations
- Pamba: Enhancing Global Interaction in Point Clouds via State Space ModelZhuoyuan Li, Yubo Ai, Jiahao Lu, Chuxin Wang et al.AAAI 2025 · 12 citations
- LinNet: Linear Network for Efficient Point Cloud Representation LearningHao Deng, Kunlei Jing, Shengmei Chen, Cheng Liu et al.NeurIPS 2024 · 12 citations
- On-the-fly Point Feature Representation for Point Clouds AnalysisJiangyi Wang, Zhongyao Cheng, Na Zhao, Jun Cheng et al.ACM MM 2024 · 9 citations
Builds on48
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- KPConv: Flexible and Deformable Convolution for Point CloudsHugues Thomas, Charles R. Qi, Jean-Emmanuel Deschaud, Beatriz Marcotegui et al.ICCV 2019 · 3,193 citations
- CCNet: Criss-Cross Attention for Semantic SegmentationZilong Huang, Xinggang Wang, Lichao Huang, Chang Huang et al.ICCV 2019 · 2,972 citations
- CvT: Introducing Convolutions to Vision TransformersHaiping Wu, Bin Xiao, Noel Codella, Mengchen Liu et al.ICCV 2021 · 2,397 citations
Related papers
- OctFormer: Octree-based Transformers for 3D Point CloudsPeng-Shuai WangSIGGRAPH 2023 · 123 citations
- PointCFormer: A Relation-Based Progressive Feature Extraction Network for Point Cloud CompletionYi Zhong, Weize Quan, Dong-Ming Yan, Jie Jiang et al.AAAI 2025 · 3 citations
- PatchFormer: An Efficient Point Transformer with Patch AttentionCheng Zhang, Haocheng Wan, Xinyi Shen, Zizhao WuCVPR 2022 · 77 citations
- FlatFormer: Flattened Window Attention for Efficient Point Cloud TransformerZhijian Liu, Xinyu Yang, Haotian Tang, Shang Yang et al.CVPR 2023
- 3D Object Detection With PointformerXuran Pan, Zhuofan Xia, Shiji Song, Li Erran Li et al.CVPR 2021
