OA-CNNs: Omni-Adaptive Sparse CNNs for 3D Semantic Segmentation
Bohao Peng, Xiaoyang Wu, Li Jiang, Yukang Chen, Hengshuang Zhao, Zhuotao Tian, Jiaya Jia
摘要
The booming of 3D recognition in the 2020s began with the introduction of point cloud transformers. They quickly overwhelmed sparse CNNs and became state-of-the-art models, especially in 3D semantic segmentation. However, sparse CNNs are still valuable networks, due to their efficiency treasure, and ease of application. In this work, we reexamine the design distinctions and test the limits of what a sparse CNN can achieve. We discover that the key credit to the performance difference is adaptivity. Specifically, we propose two key components, i.e., adaptive receptive fields (spatially) and adaptive relation, to bridge the gap. This exploration led to the creation of Omni-Adaptive 3D CNNs (OA-CNNs), a family of networks that integrates a lightweight module to greatly enhance the adaptivity of sparse CNNs at minimal computational cost. Without any self-attention modules, OA-CNNs favorably surpass point transformers in terms of accuracy in both indoor and outdoor scenes, with much less latency and memory cost. Notably, it achieves 76.1%, 78.9%, and 70.6% mIoU on ScanNet v2, nuScenes, and SemanticKITTI validation benchmarks respectively, while maintaining at most 5× better speed than transformer counterparts. This revelation highlights the potential of pure sparse CNNs to outperform transformer-related networks. Our code is built upon Pointcept [9], which is available at here<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup><sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup>https://github.com/Pointcept/Pointcept.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper33
- Towards Large-Scale 3D Representation Learning with Multi-Dataset Point Prompt TrainingXiaoyang Wu, Zhuotao Tian, Xin Wen, Bohao Peng 等CVPR 2024 · 被引用 39 次
- LitePT: Lighter Yet Stronger Point TransformerYuanwen Yue, Damien Robert, Jianyuan Wang, Sunghwan Hong 等CVPR 2026 · 被引用 25 次
- Less Is More, but Where? Dynamic Token Compression via LLM-Guided Keyframe PriorYulin Li, Haokun Gui, Ziyang Fan, Junjie Wang 等NeurIPS 2025 · 被引用 18 次
- FlashVID: Efficient Video Large Language Models via Training-free Tree-based Spatiotemporal Token MergingZiyang Fan, Keyu Chen, Ruilong Xing, Yulin Li 等ICLR 2026 · 被引用 15 次
- LinNet: Linear Network for Efficient Point Cloud Representation LearningHao Deng, Kunlei Jing, Shengmei Chen, Cheng Liu 等NeurIPS 2024 · 被引用 12 次
它引用的顶会 Paper32
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- KPConv: Flexible and Deformable Convolution for Point CloudsHugues Thomas, Charles R. Qi, Jean-Emmanuel Deschaud, Beatriz Marcotegui 等ICCV 2019 · 被引用 3,193 次
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel 等ICCV 2019 · 被引用 2,345 次
相关 Paper
- OctFormer: Octree-based Transformers for 3D Point CloudsPeng-Shuai WangSIGGRAPH 2023 · 被引用 123 次
- Interpolation-Aware Padding for 3D Sparse Convolutional Neural NetworksYu-Qi Yang, Peng-Shuai Wang, Yang LiuICCV 2021 · 被引用 4 次
- Focal Sparse Convolutional Networks for 3D Object DetectionYukang Chen, Yanwei Li, Xiangyu Zhang, Jian Sun 等CVPR 2022 · 被引用 293 次
- Point Transformer V2: Grouped Vector Attention and Partition-based PoolingXiaoyang Wu, Yixing Lao, Li Jiang, Xihui Liu 等NeurIPS 2022 · 被引用 924 次
- SVT-Net: Super Light-Weight Sparse Voxel Transformer for Large Scale Place RecognitionZhaoxin Fan, Zhenbo Song, Hongyan Liu, Zhiwu Lu 等AAAI 2022 · 被引用 95 次
