Point Cloud Self-Supervised Learning via 3D to Multi-View Masked Learner
Zhimin Chen, Xuewei Chen, Xiao Guo, Yingwei Li, Longlong Jing, Liang Yang, Bing Li
摘要
Recently, multi-modal masked autoencoders (MAE) has been introduced in 3D self-supervised learning, offering enhanced feature learning by leveraging both 2D and 3D data to capture richer cross-modal representations. However, these approaches have two limitations: (1) they inefficiently require both 2D and 3D modalities as inputs, even though the inherent multi-view properties of 3D point clouds already contain modality. (2) input modality causes the reconstruction learning to unnecessarily rely on visible 2D information, hindering 3D geometric representation learning. To address these challenges, we propose a 3D to Multi-View Learner (Multi-View ML) that only utilizes 3D modalities as inputs and effectively capture rich spatial information in 3D point clouds. Specifically, we first project 3D point clouds to multi-view 2D images at the feature level based on 3D-based pose. Then, we introduce two components: (1) a 3D to multi-view autoencoder that reconstructs point clouds and multi-view images from 3D and projected 2D features; (2) a multi-scale multi-head (MSMH) attention mechanism that facilitates local-global information interactions in each decoder transformer block through attention heads at various scales. Additionally, a novel twostage self-training strategy is proposed to align and representations. Our method outperforms state-of-the-art counterparts across various downstream tasks, including 3D classification, part segmentation, and object detection.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- TokenUnify: Scaling Up Autoregressive Pretraining for Neuron SegmentationYinda Chen, Haoyuan Shi, Xiaoyu Liu, Te Shi 等ICCV 2025 · 被引用 1 次
- Dual-Calibration Multi-View Clustering via Compact Anchor LearningHuibing Wang, Yuemeng Huang, Yawei Chen, Jiaxin Yang 等ICML 2026
它引用的顶会 Paper30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- CenterNet: Keypoint Triplets for Object DetectionKaiwen Duan, Song Bai, Lingxi Xie, Honggang Qi 等ICCV 2019 · 被引用 3,348 次
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 被引用 1,467 次
- Revisiting Point Cloud Classification: A New Benchmark Dataset and Classification Model on Real-World DataMikaela Angelina Uy, Quang-Hieu Pham, Binh-Son Hua, Duc Thanh Nguyen 等ICCV 2019 · 被引用 1,003 次
- Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point ModelingXumin Yu, Lulu Tang, Yongming Rao, Tiejun Huang 等CVPR 2022 · 被引用 684 次
相关 Paper
- Reliable-View 2D-3D Key-Part Aligned Transformer with Reinforced Masking for 3D Point Cloud UnderstandingXianglong Jin, Zheng Wang, Rong Wang, Feiping NieAAAI 2026
- Point-M2AE: Multi-scale Masked Autoencoders for Hierarchical Point Cloud Pre-trainingRenrui Zhang, Ziyu Guo, Peng Gao, Rongyao Fang 等NeurIPS 2022 · 被引用 445 次
- PiMAE: Point Cloud and Image Interactive Masked Autoencoders for 3D Object DetectionAnthony Chen, Kevin Zhang, Renrui Zhang, Zihan Wang 等CVPR 2023
- Learning 3D Representations from 2D Pre-Trained Models via Image-to-Point Masked AutoencodersRenrui Zhang, Liuhui Wang, Yu Qiao, Peng Gao 等CVPR 2023
- 2D-3D Interlaced Transformer for Point Cloud Segmentation with Scene-Level SupervisionCheng-Kun Yang, Min-Hung Chen, Yung-Yu Chuang, Yen-Yu LinICCV 2023 · 被引用 30 次
