Towards More Diverse and Challenging Pre-Training for Point Cloud Learning: Self-Supervised Cross Reconstruction with Decoupled Views
Xiangdong Zhang, Shaofeng Zhang, Junchi Yan
Abstract
Point cloud learning, especially in a self-supervised way without manual labels, has gained growing attention in both vision and learning communities due to its potential utility in a wide range of applications. Most existing generative approaches for point cloud self-supervised learning focus on recovering masked points from visible ones within a single view. Recognizing that a two-view pre-training paradigm inherently introduces greater diversity and variance, it may thus enable more challenging and informative pre-training. Inspired by this, we explore the potential of two-view learning in this domain. In this paper, we propose Point-PQAE, a cross-reconstruction generative paradigm that first generates two decoupled point clouds/views and then reconstructs one from the other. To achieve this goal, we develop a crop mechanism for point cloud view generation for the first time and further propose a novel positional encoding to represent the 3D relative position between the two decoupled views. The cross-reconstruction significantly increases the difficulty of pre-training compared to self-reconstruction, which enables our method to surpass previous single-modal self-reconstruction methods in 3D self-supervised learning. Specifically, it outperforms the self-reconstruction baseline (Point-MAE) by 6.5%, 7.0%, and 6.7% in three variants of ScanObjectNN with the Mlp-Linear evaluation protocol. The code is available at https://github.com/aHapBean/Point-PQAE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 30f92f6b-7d58-4d93-be71-8bd64f4e49e3Cited by top-tier papers3
- PointCSP: Cross-Sample Semantic Propagation and Stability Preservation in Self-Supervised Point Cloud LearningXinxing Yu, Ajian Liu, Sunyuan Qiang, Hui Ma et al.CVPR 2026 · 1 citation
- PointAlign: Feature-Level Alignment Regularization for 3D Vision-Language ModelsYuanhao Su, Shaofeng Zhang, Xiaosong Jia, Qi FanCVPR 2026 · 1 citation
- PointCHR: Point Cloud Analysis via Curvature-Aware Hyperbolic RectificationXinxing Yu, Liying Yang, Hao Mo, Hui Ma et al.ICML 2026
Builds on46
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 1,467 citations
Related papers
- PCP-MAE: Learning to Predict Centers for Point Masked AutoencodersXiangdong Zhang, Shaofeng Zhang, Junchi YanNeurIPS 2024 · 44 citations
- Regress Before Construct: Regress Autoencoder for Point Cloud Self-supervised LearningYang Liu, Chen Chen, Can Wang, Xulin King et al.ACM MM 2023 · 13 citations
- Point-MaDi: Masked Autoencoding with Diffusion for Point Cloud Pre-trainingXiaoyang Xiao, Runzhao Yao, Zhiqiang Tian, Shaoyi DuNeurIPS 2025 · 4 citations
- Take-A-Photo: 3D-to-2D Generative Pre-training of Point Cloud ModelsZiyi Wang, Xumin Yu, Yongming Rao, Jie Zhou et al.ICCV 2023 · 34 citations
- PiMAE: Point Cloud and Image Interactive Masked Autoencoders for 3D Object DetectionAnthony Chen, Kevin Zhang, Renrui Zhang, Zihan Wang et al.CVPR 2023
