Multimodal Robust Prompt Distillation for 3D Point Cloud Models
Xiang Gu, Liming Lu, Xu Zheng, Anan Du, Yongbin Zhou, Shuchao Pang
Abstract
Adversarial attacks pose a significant threat to learning-based 3D point cloud models, critically undermining their reliability in security-sensitive applications. Existing defense methods often suffer from (1) high computational overhead and (2) poor generalization ability across diverse attack types. To bridge these gaps, we propose a novel yet efficient teacher-student framework, namely Multimodal Robust Prompt Distillation (MRPD) for distilling robust 3D point cloud model. It learns lightweight prompts by aligning student point cloud model's features with robust embeddings from three distinct teachers: a vision model processing depth projections, a high-performance 3D model, and a text encoder. To ensure a reliable knowledge transfer, this distillation is guided by a confidence-gated mechanism which dynamically balances the contribution of all input modalities. Notably, since the distillation is all during the training stage, there is no additional computational cost at inference. Extensive experiments demonstrate that MRPD substantially outperforms state-of-the-art defense methods against a wide range of white-box and black-box attacks, while even achieving better performance on clean data. Our work presents a new, practical paradigm for building robust 3D vision systems by efficiently harnessing multimodal knowledge.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- PointCLIP V2: Prompting CLIP and GPT for Powerful 3D Open-world LearningXiangyang Zhu, Renrui Zhang, Bowei He, Ziyu Guo et al.ICCV 2023 · 248 citations
- CLIP2Point: Transfer CLIP to Point Cloud Classification with Image-Depth Pre-TrainingTianyu Huang, Bowen Dong, Yunhan Yang, Xiaoshui Huang et al.ICCV 2023 · 220 citations
- Uni3D: Exploring Unified 3D Representation at ScaleJunsheng Zhou, Jinsheng Wang, Baorui Ma, Yu-Shen Liu et al.ICLR 2024 · 207 citations
- DUP-Net: Denoiser and Upsampler Network for 3D Adversarial Point Clouds DefenseHang Zhou, Kejiang Chen, Weiming Zhang, Han Fang et al.ICCV 2019 · 206 citations
Related papers
- A Critical Revisit of Adversarial Robustness in 3D Point Cloud Recognition with Diffusion-Driven PurificationJiachen Sun, Jiongxiao Wang, Weili Nie, Zhiding Yu et al.ICML 2023 · 24 citations
- Ada3Diff: Defending against 3D Adversarial Point Clouds via Adaptive DiffusionKui Zhang, Hang Zhou, Jie Zhang, Qidong Huang et al.ACM MM 2023 · 14 citations
- Let Images Give You More: Point Cloud Cross-Modal Training for Shape AnalysisXu Yan, Heshen Zhan, Chaoda Zheng, Jiantao Gao et al.NeurIPS 2022 · 49 citations
- MemDistill: Distilling LiDAR Knowledge into Memory for Camera-Only 3D Object DetectionDonghyeon Kwon, Youngseok Yoon, Hyeongseok Son, Suha KwakICCV 2025 · 1 citation
- PointDistiller: Structured Knowledge Distillation Towards Efficient and Compact 3D DetectionLinfeng Zhang, Runpei Dong, Hung-Shuo Tai, Kaisheng MaCVPR 2023
