Exploring Vision Semantic Prompt for Efficient Point Cloud Understanding
Yixin Zha, Chuxin Wang, Wenfei Yang, Tianzhu Zhang, Feng Wu
Abstract
A series of pretrained models have demonstrated promising results in point cloud understanding tasks and are widely applied to downstream tasks through fine-tuning. However, full fine-tuning leads to the forgetting of pretrained knowledge and substantial storage costs on edge devices. To address these issues, Parameter-Efficient Transfer Learning (PETL) methods have been proposed. According to our analysis, we find that existing 3D PETL methods cannot adequately align with semantic relationships of features required by downstream tasks, resulting in suboptimal performance. To ensure parameter efficiency while introducing rich semantic cues, we propose a novel fine-tuning paradigm for 3D pretrained models. We utilize frozen 2D pretrained models to provide vision semantic prompts and design a new Hybrid Attention Adapter to efficiently fuse 2D semantic cues into 3D representations with minimal trainable parameters(1.8M). Extensive experiments conducted on datasets including ScanOb-jectNN, ModelNet40, and ShapeNetPart demonstrate the effectiveness of our proposed paradigm. In particular, our method achieves 95.6% accuracy on ModelNet40 and attains 90.09% performance on the most challenging classification split ScanObjectNN(PB-T50-RS). Layer 0 Layer 8 Pretrained Full FT PEFT Ours Layer 4 (a) The feature distributions across different layers of various methods. 3d-to-2d Proj. Local Ambiguity Global Ambiguity (c) The ambiguity between 3D and 2D. (b) Performences on ScanObjectNN(Hardest).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- FS-I2P: A Hierarchical Focus–Sweep Registration Network with Dynamically Allocated DepthZhixin Cheng, Yujia Chen, Xujing Tao, Bohao Liao et al.ICML 2026 · 2 citations
- Rethinking 2D-3D Registration: A Novel Network for High-Value Zone Selection and Representation Consistency AlignmentZhixin Cheng, Bohao Liao, Jiacheng Deng, Xiaotian Yin et al.CVPR 2026 · 2 citations
- PointCHR: Point Cloud Analysis via Curvature-Aware Hyperbolic RectificationXinxing Yu, Liying Yang, Hao Mo, Hui Ma et al.ICML 2026
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- AdaptFormer: Adapting Vision Transformers for Scalable Visual RecognitionShoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang et al.NeurIPS 2022 · 1,291 citations
- Towards a Unified View of Parameter-Efficient Transfer LearningJunxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick et al.ICLR 2022 · 1,182 citations
Related papers
- Positional Prompt Tuning for Efficient 3D Representation LearningShaochen Zhang, Zekun Qi, Runpei Dong, Xiuxiu Bai et al.ACM MM 2025 · 2 citations
- Point-PEFT: Parameter-Efficient Fine-Tuning for 3D Pre-trained ModelsYiwen Tang, Ray Zhang, Zoey Guo, Xianzheng Ma et al.AAAI 2024
- GAPrompt: Geometry-Aware Point Cloud Prompt for 3D Vision ModelZixiang Ai, Zichen Liu, Yuanhang Lei, Zhenyu Cui et al.ICML 2025
- P2P: Tuning Pre-trained Image Models for Point Cloud Analysis with Point-to-Pixel PromptingZiyi Wang, Xumin Yu, Yongming Rao, Jie Zhou et al.NeurIPS 2022 · 121 citations
- On Geometry-Enhanced Parameter-Efficient Fine-Tuning for 3D Scene SegmentationLiyao Tang, Zhe Chen, Dacheng TaoNeurIPS 2025 · 5 citations
