PointRWKV: Efficient RWKV-Like Model for Hierarchical Point Cloud Learning
Qingdong He, Jiangning Zhang, Jinlong Peng, Haoyang He, Xiangtai Li, Yabiao Wang, Chengjie Wang
摘要
Transformers have revolutionized the point cloud learning task, but the quadratic complexity hinders its extension to long sequences. This puts a burden on limited computational resources. The recent advent of RWKV, a fresh breed of deep sequence models, has shown immense potential for sequence modeling in NLP tasks. In this work, we present PointRWKV, a new model of linear complexity derived from the RWKV model in the NLP field with the necessary adaptation for 3D point cloud learning tasks. Specifically, taking the embedded point patches as input, we first propose to explore the global processing capabilities within PointRWKV blocks using modified multi-headed matrix-valued states and a dynamic attention recurrence mechanism. To extract local geometric features simultaneously, we design a parallel branch to encode the point cloud efficiently in a fixed radius near-neighbors graph with a graph stabilizer. Furthermore, we design PointR-WKV as a multi-scale framework for hierarchical feature learning of 3D point clouds, facilitating various downstream tasks. Extensive experiments on different point cloud learning tasks show our proposed PointRWKV outperforms the transformer-and mamba-based counterparts, while significantly saving about 42% FLOPs, demonstrating the potential option for constructing foundational 3D models. Project page: https://hithqd.github. io/projects/PointRWKV/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- RWKV-CLIP: A Robust Vision-Language Representation LearnerTiancheng Gu, Kaicheng Yang, Xiang An, Ziyong Feng 等EMNLP 2024 · 被引用 11 次
- CLIP-GS: Unifying Vision-Language Representation with 3D Gaussian SplattingSiyu Jiao, Haoye Dong, Yuyang Yin, Zequn Jie 等ICCV 2025 · 被引用 4 次
- Positional Prompt Tuning for Efficient 3D Representation LearningShaochen Zhang, Zekun Qi, Runpei Dong, Xiuxiu Bai 等ACM MM 2025 · 被引用 2 次
- FractalCloud: A Fractal-Inspired Architecture for Efficient Large-Scale Point Cloud ProcessingYuzhe Fu, Changchun Zhou, Hancheng Ye, Bowen Duan 等HPCA 2026 · 被引用 1 次
- ProConMV: Provenance-Enabled Conceptual Framework for Interpretable Multi-View Diabetic Retinopathy DiagnosisXiaoling Luo, Shuo Yang, Qihao Xu, Jiansong Zhang 等ICML 2026
它引用的顶会 Paper30
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu 等NeurIPS 2024 · 被引用 3,199 次
- KPConv: Flexible and Deformable Convolution for Point CloudsHugues Thomas, Charles R. Qi, Jean-Emmanuel Deschaud, Beatriz Marcotegui 等ICCV 2019 · 被引用 3,193 次
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang 等ICML 2024 · 被引用 1,725 次
相关 Paper
- RWKV3D: An RWKV-Based Model with Multiple Training Strategies for Point Cloud AnalysisChenglong Sun, Shijie Pang, Yuzheng Wang, Lizhe QiACM MM 2025
- PointDGRWKV: Generalizing RWKV-like Architecture to Unseen Domains for Point Cloud ClassificationHao Yang, Qianyu Zhou, Haijia Sun, Xiangtai Li 等AAAI 2026
- PatchFormer: An Efficient Point Transformer with Patch AttentionCheng Zhang, Haocheng Wan, Xinyi Shen, Zizhao WuCVPR 2022 · 被引用 77 次
- PointMamba: A Simple State Space Model for Point Cloud AnalysisDingkang Liang, Xin Zhou, Wei Xu, Xingkui Zhu 等NeurIPS 2024 · 被引用 380 次
- Mamba3D: Enhancing Local Features for 3D Point Cloud Analysis via State Space ModelXu Han, Yuan Tang, Zhaoxuan Wang, Xianzhi LiACM MM 2024 · 被引用 86 次
