RI-MAE: Rotation-Invariant Masked AutoEncoders for Self-Supervised Point Cloud Representation Learning
Kunming Su, Qiuxia Wu, Panpan Cai, Xiaogang Zhu, Xuequan Lu, Zhiyong Wang, Kun Hu
Abstract
Masked point modeling methods have recently achieved great success in self-supervised learning for point cloud data. However, these methods are sensitive to rotations and often exhibit sharp performance drops when encountering rotational variations. In this paper, we propose a novel Rotation-Invariant Masked AutoEncoders (RI-MAE) to address two major challenges: 1) achieving rotation-invariant latent representations, and 2) facilitating self-supervised reconstruction in a rotation-invariant manner. For the first challenge, we introduce RI-Transformer, which features disentangled geometry content, rotation-invariant relative orientation and position embedding mechanisms for constructing rotationinvariant point cloud latent space. For the second challenge, a novel dual-branch student-teacher architecture is devised. It enables the self-supervised learning via the reconstruction of masked patches within the learned rotation-invariant latent space. Each branch is based on an RI-Transformer, and they are connected with an additional RI-Transformer predictor. The teacher encodes all point patches, while the student solely encodes unmasked ones. Finally, the predictor predicts the latent features of the masked patches using the output latent embeddings from the student, supervised by the outputs from the teacher. Extensive experiments demonstrate that our method is robust to rotations, achieving the state-of-the-art performance on various downstream tasks. Our code is available at https://github.com/kunmingsu07/RI-MAE .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ff2dd1db-221a-45f4-87cf-18a8a4672990Cited by top-tier papers9
- Rotary Masked Autoencoders are Versatile LearnersUros Zivanovic, Serafina Di Gioia, Andre Scaffidi, Martín de los Rios et al.NeurIPS 2025 · 4 citations
- PUMPS: Skeleton-Agnostic Point-Based Universal Motion Pre-Training for Synthesis in Human Motion TasksClinton Ansun Mo, Kun Hu, Chengjiang Long, Dong Yuan et al.ICCV 2025 · 2 citations
- PointMC: Multi-view Consistent Encoding and Center-Global Feature Fusion for Point Clouds UnderstandingXinxing Yu, Ajian Liu, Sunyuan Qiang, Yuzhong Wang et al.AAAI 2026 · 1 citation
- PhenoYieldNet: Learning Crop-Aware Phenological Responses for Multi-Crop Yield PredictionYu Luo, Xiaogang Zhu, Shan Zeng, Wei Xiang et al.CVPR 2026 · 1 citation
- Reliable-View 2D-3D Key-Part Aligned Transformer with Reinforced Masking for 3D Point Cloud UnderstandingXianglong Jin, Zheng Wang, Rong Wang, Feiping NieAAAI 2026
Builds on17
- Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point ModelingXumin Yu, Lulu Tang, Yongming Rao, Tiejun Huang et al.CVPR 2022 · 684 citations
- Rethinking and Improving Relative Position Encoding for Vision TransformerKan Wu, Houwen Peng, Minghao Chen, Jianlong Fu et al.ICCV 2021 · 427 citations
- Unsupervised Point Cloud Pre-training via Occlusion CompletionHanchen Wang, Qi Liu, Xiangyu Yue, Joan Lasenby et al.ICCV 2021 · 323 citations
- CrossPoint: Self-Supervised Cross-Modal Contrastive Learning for 3D Point Cloud UnderstandingMohamed Afham, Isuru Dissanayake, Dinithi Dissanayake, Amaya Dharmasiri et al.CVPR 2022 · 286 citations
- Contrast with Reconstruct: Contrastive 3D Representation Learning Guided by Generative PretrainingZekun Qi, Runpei Dong, Guofan Fan, Zheng Ge et al.ICML 2023 · 209 citations
Related papers
- Regress Before Construct: Regress Autoencoder for Point Cloud Self-supervised LearningYang Liu, Chen Chen, Can Wang, Xulin King et al.ACM MM 2023 · 13 citations
- 3D Feature Prediction for Masked-AutoEncoder-Based Point Cloud PretrainingSiming Yan, Yuqi Yang, Yu-Xiao Guo, Hao Pan et al.ICLR 2024 · 21 citations
- Point Cloud Self-Supervised Learning via 3D to Multi-View Masked LearnerZhimin Chen, Xuewei Chen, Xiao Guo, Yingwei Li et al.ICCV 2025 · 1 citation
- Point-M2AE: Multi-scale Masked Autoencoders for Hierarchical Point Cloud Pre-trainingRenrui Zhang, Ziyu Guo, Peng Gao, Rongyao Fang et al.NeurIPS 2022 · 445 citations
- PCP-MAE: Learning to Predict Centers for Point Masked AutoencodersXiangdong Zhang, Shaofeng Zhang, Junchi YanNeurIPS 2024 · 44 citations
