CodedVTR: Codebook-based Sparse Voxel Transformer with Geometric Guidance
Tianchen Zhao, Niansong Zhang, Xuefei Ning, He Wang, Li Yi, Yu Wang
Abstract
Transformers have gained much attention by outperforming convolutional neural networks in many 2D vision tasks. However, they are known to have generalization problems and rely on massive-scale pre-training and sophisticated training techniques. When applying to 3D tasks, the irregular data structure and limited data scale add to the difficulty of transformer's application. We propose Cod-edVTR (Codebook-based Voxel TRansformer), which improves data efficiency and generalization ability for 3D sparse voxel transformers. On the one hand, we propose the codebook-based attention that projects an attention space into its subspace represented by the combination of "prototypes" in a learnable codebook. It regularizes attention learning and improves generalization. On the other hand, we propose geometry-aware self-attention that utilizes geometric information (geometric pattern, density) to guide attention learning. CodedVTR could be embedded into existing sparse convolution-based methods, and bring consistent performance improvements for indoor and outdoor 3D semantic segmentation tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 185d1b95-debe-428b-a260-45ef4a8e7aeeCited by top-tier papers5
- Ada3D : Exploiting the Spatial Redundancy with Adaptive Inference for Efficient 3D Object DetectionTianchen Zhao, Xuefei Ning, Ke Hong, Zhongyuan Qiu et al.ICCV 2023 · 23 citations
- DQS3D: Densely-matched Quantization-aware Semi-supervised 3D DetectionHuan-ang Gao, Beiwen Tian, Pengfei Li, Hao Zhao et al.ICCV 2023 · 21 citations
- SEFormer: Structure Embedding Transformer for 3D Object DetectionXiaoyu Feng, Heming Du, Hehe Fan, Yueqi Duan et al.AAAI 2023 · 15 citations
- Decompose Novel into Known: Part Concept Learning For 3D Novel Class DiscoveryTingyu Weng, Jun Xiao, Haiyong JiangNeurIPS 2023 · 6 citations
- Activating Sparse Part Concepts for 3D Class Incremental LearningZhenya Tian, Jun Xiao, Lupeng Liu, Haiyong JiangCVPR 2025
Builds on17
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- KPConv: Flexible and Deformable Convolution for Point CloudsHugues Thomas, Charles R. Qi, Jean-Emmanuel Deschaud, Beatriz Marcotegui et al.ICCV 2019 · 3,193 citations
- CvT: Introducing Convolutions to Vision TransformersHaiping Wu, Bin Xiao, Noel Codella, Mengchen Liu et al.ICCV 2021 · 2,397 citations
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel et al.ICCV 2019 · 2,345 citations
Related papers
- GTA: A Geometry-Aware Attention Mechanism for Multi-View TransformersTakeru Miyato, Bernhard Jaeger, Max Welling, Andreas GeigerICLR 2024 · 51 citations
- Fast Point TransformerChunghyun Park, Yoonwoo Jeong, Minsu Cho, Jaesik ParkCVPR 2022
- Voxel Set Transformer: A Set-to-Set Approach to 3D Object Detection from Point CloudsChenhang He, Ruihuang Li, Shuai Li, Lei ZhangCVPR 2022 · 217 citations
- Voxel Transformer for 3D Object DetectionJiageng Mao, Yujing Xue, Minzhe Niu, Haoyue Bai et al.ICCV 2021 · 535 citations
- OctFormer: Octree-based Transformers for 3D Point CloudsPeng-Shuai WangSIGGRAPH 2023 · 123 citations
