RWKV3D: An RWKV-Based Model with Multiple Training Strategies for Point Cloud Analysis
Chenglong Sun, Shijie Pang, Yuzheng Wang, Lizhe Qi
Abstract
Transformer-based models have achieved dominance in point cloud analysis, yet their quadratic computational complexity remains a fundamental limitation for practical applications. Recently, RWKV has emerged as a promising alternative for sequence modeling due to its linear computational complexity. However, it has yet to be effectively adapted to handle the unordered and sparse nature of point cloud data. In this paper, we propose RWKV3D, an innovative and computational framework tailored for point cloud analysis, which is adaptable to three training strategies: training from scratch, single-modal pre-training, and cross-modal pre-training. First, we replace the MLP layer with an advanced Local Feature Mixer (LFM), which not only enhances fine-grained feature extraction but also reduces the number of parameters. Second, we introduce a Bidirectional Multi-head Shift (BMS) mechanism to expand the receptive field, effectively capturing richer contextual information. Additionally, to enhance high-level feature processing, we strategically incorporate a Multi-head Self-Attention (MSA) block before the first RWKV3D block. Experimental results demonstrate that RWKV3D outperforms Transformer-based and Mamba-based methods while maintaining lower parameter counts and computational costs. Notably, it achieves several state-of-the-art results, including overall accuracies of 95.3% (training from scratch) and 95.9% (cross-modal pre-training) on the ModelNet40 dataset, as well as 95.28% (single-modal pre-training) on the ScanObjectNN (PB_T50_RS) dataset. These results underscore the superior efficacy of the RWKV architecture in 3D vision tasks and highlight its potential for broader multimodal learning scenarios.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 1fce5538-fbe0-4a09-97b6-ed7d2605504eRelated papers
- Mamba3D: Enhancing Local Features for 3D Point Cloud Analysis via State Space ModelXu Han, Yuan Tang, Zhaoxuan Wang, Xianzhi LiACM MM 2024 · 86 citations
- PointRWKV: Efficient RWKV-Like Model for Hierarchical Point Cloud LearningQingdong He, Jiangning Zhang, Jinlong Peng, Haoyang He et al.AAAI 2025 · 41 citations
- PointDGRWKV: Generalizing RWKV-like Architecture to Unseen Domains for Point Cloud ClassificationHao Yang, Qianyu Zhou, Haijia Sun, Xiangtai Li et al.AAAI 2026
- LCM: Locally Constrained Compact Point Cloud Model for Masked Point ModelingYaohua Zha, Naiqi Li, Yanzi Wang, Tao Dai et al.NeurIPS 2024 · 25 citations
- LION: Linear Group RNN for 3D Object Detection in Point CloudsZhe Liu, Jinghua Hou, Xinyu Wang, Xiaoqing Ye et al.NeurIPS 2024 · 84 citations
