CodedVTR: Codebook-based Sparse Voxel Transformer with Geometric Guidance
Tianchen Zhao, Niansong Zhang, Xuefei Ning, He Wang, Li Yi, Yu Wang
摘要
Transformers have gained much attention by outperforming convolutional neural networks in many 2D vision tasks. However, they are known to have generalization problems and rely on massive-scale pre-training and sophisticated training techniques. When applying to 3D tasks, the irregular data structure and limited data scale add to the difficulty of transformer's application. We propose Cod-edVTR (Codebook-based Voxel TRansformer), which improves data efficiency and generalization ability for 3D sparse voxel transformers. On the one hand, we propose the codebook-based attention that projects an attention space into its subspace represented by the combination of "prototypes" in a learnable codebook. It regularizes attention learning and improves generalization. On the other hand, we propose geometry-aware self-attention that utilizes geometric information (geometric pattern, density) to guide attention learning. CodedVTR could be embedded into existing sparse convolution-based methods, and bring consistent performance improvements for indoor and outdoor 3D semantic segmentation tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Ada3D : Exploiting the Spatial Redundancy with Adaptive Inference for Efficient 3D Object DetectionTianchen Zhao, Xuefei Ning, Ke Hong, Zhongyuan Qiu 等ICCV 2023 · 被引用 23 次
- DQS3D: Densely-matched Quantization-aware Semi-supervised 3D DetectionHuan-ang Gao, Beiwen Tian, Pengfei Li, Hao Zhao 等ICCV 2023 · 被引用 21 次
- SEFormer: Structure Embedding Transformer for 3D Object DetectionXiaoyu Feng, Heming Du, Hehe Fan, Yueqi Duan 等AAAI 2023 · 被引用 15 次
- Decompose Novel into Known: Part Concept Learning For 3D Novel Class DiscoveryTingyu Weng, Jun Xiao, Haiyong JiangNeurIPS 2023 · 被引用 6 次
- Activating Sparse Part Concepts for 3D Class Incremental LearningZhenya Tian, Jun Xiao, Lupeng Liu, Haiyong JiangCVPR 2025
它引用的顶会 Paper17
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- KPConv: Flexible and Deformable Convolution for Point CloudsHugues Thomas, Charles R. Qi, Jean-Emmanuel Deschaud, Beatriz Marcotegui 等ICCV 2019 · 被引用 3,193 次
- CvT: Introducing Convolutions to Vision TransformersHaiping Wu, Bin Xiao, Noel Codella, Mengchen Liu 等ICCV 2021 · 被引用 2,397 次
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel 等ICCV 2019 · 被引用 2,345 次
相关 Paper
- GTA: A Geometry-Aware Attention Mechanism for Multi-View TransformersTakeru Miyato, Bernhard Jaeger, Max Welling, Andreas GeigerICLR 2024 · 被引用 51 次
- Fast Point TransformerChunghyun Park, Yoonwoo Jeong, Minsu Cho, Jaesik ParkCVPR 2022
- Voxel Set Transformer: A Set-to-Set Approach to 3D Object Detection from Point CloudsChenhang He, Ruihuang Li, Shuai Li, Lei ZhangCVPR 2022 · 被引用 217 次
- Voxel Transformer for 3D Object DetectionJiageng Mao, Yujing Xue, Minzhe Niu, Haoyue Bai 等ICCV 2021 · 被引用 535 次
- OctFormer: Octree-based Transformers for 3D Point CloudsPeng-Shuai WangSIGGRAPH 2023 · 被引用 123 次
