Take-A-Photo: 3D-to-2D Generative Pre-training of Point Cloud Models
Ziyi Wang, Xumin Yu, Yongming Rao, Jie Zhou, Jiwen Lu
摘要
With the overwhelming trend of mask image modeling led by MAE, generative pre-training has shown a remarkable potential to boost the performance of fundamental models in 2D vision. However, in 3D vision, the over-reliance on Transformer-based backbones and the unordered nature of point clouds have restricted the further development of generative pre-training. In this paper, we propose a novel 3D-to-2D generative pre-training method that is adaptable to any point cloud model. We propose to generate view images from different instructed poses via the cross-attention mechanism as the pre-training scheme. Generating view images has more precise supervision than its point cloud counterpart, thus assisting 3D backbones to have a finer comprehension of the geometrical structure and stereoscopic relations of the point cloud. Experimental results have proved the superiority of our proposed 3D-to-2D generative pre-training over previous pre-training methods. Our method is also effective in boosting the performance of architecture-oriented approaches, achieving state-of-the-art performance when fine-tuning on ScanObjectNN classification and ShapeNet-Part segmentation tasks. Code is available at https://github.com/wangzy22/TakeAPhoto.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- PCP-MAE: Learning to Predict Centers for Point Masked AutoencodersXiangdong Zhang, Shaofeng Zhang, Junchi YanNeurIPS 2024 · 被引用 44 次
- Point Cloud Pre-Training with Diffusion ModelsXiao Zheng, Xiaoshui Huang, Guofeng Mei, Yuenan Hou 等CVPR 2024 · 被引用 33 次
- Frozen CLIP Transformer Is an Efficient Point Cloud EncoderXiaoshui Huang, Zhou Huang, Sheng Li, Wentao Qu 等AAAI 2024 · 被引用 30 次
- RepKPU: Point Cloud Upsampling with Kernel Point Representation and DeformationYi Rong, Haoran Zhou, Kang Xia, Cheng Mei 等CVPR 2024 · 被引用 22 次
- Towards More Diverse and Challenging Pre-Training for Point Cloud Learning: Self-Supervised Cross Reconstruction with Decoupled ViewsXiangdong Zhang, Shaofeng Zhang, Junchi YanICCV 2025 · 被引用 4 次
它引用的顶会 Paper25
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi 等NeurIPS 2020 · 被引用 2,611 次
- PointNeXt: Revisiting PointNet++ with Improved Training and Scaling StrategiesGuocheng Qian, Yuchen Li, Houwen Peng, Jinjie Mai 等NeurIPS 2022 · 被引用 1,270 次
- Revisiting Point Cloud Classification: A New Benchmark Dataset and Classification Model on Real-World DataMikaela Angelina Uy, Quang-Hieu Pham, Binh-Son Hua, Duc Thanh Nguyen 等ICCV 2019 · 被引用 1,003 次
相关 Paper
- P2P: Tuning Pre-trained Image Models for Point Cloud Analysis with Point-to-Pixel PromptingZiyi Wang, Xumin Yu, Yongming Rao, Jie Zhou 等NeurIPS 2022 · 被引用 121 次
- Learning 3D Representations from 2D Pre-Trained Models via Image-to-Point Masked AutoencodersRenrui Zhang, Liuhui Wang, Yu Qiao, Peng Gao 等CVPR 2023
- Point-M2AE: Multi-scale Masked Autoencoders for Hierarchical Point Cloud Pre-trainingRenrui Zhang, Ziyu Guo, Peng Gao, Rongyao Fang 等NeurIPS 2022 · 被引用 445 次
- Point-MaDi: Masked Autoencoding with Diffusion for Point Cloud Pre-trainingXiaoyang Xiao, Runzhao Yao, Zhiqiang Tian, Shaoyi DuNeurIPS 2025 · 被引用 4 次
- Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point ModelingXumin Yu, Lulu Tang, Yongming Rao, Tiejun Huang 等CVPR 2022 · 被引用 684 次
