Let Images Give You More: Point Cloud Cross-Modal Training for Shape Analysis
Xu Yan, Heshen Zhan, Chaoda Zheng, Jiantao Gao, Ruimao Zhang, Shuguang Cui, Zhen Li
摘要
Although recent point cloud analysis achieves impressive progress, the paradigm of representation learning from a single modality gradually meets its bottleneck. In this work, we take a step towards more discriminative 3D point cloud representation by fully taking advantages of images which inherently contain richer appearance information, e.g., texture, color, and shade. Specifically, this paper introduces a simple but effective point cloud cross-modality training (PointCMT) strategy, which utilizes view-images, i.e., rendered or projected 2D images of the 3D object, to boost point cloud analysis. In practice, to effectively acquire auxiliary knowledge from view images, we develop a teacher-student framework and formulate the cross modal learning as a knowledge distillation problem. PointCMT eliminates the distribution discrepancy between different modalities through novel feature and classifier enhancement criteria and avoids potential negative transfer effectively. Note that PointCMT effectively improves the point-only representation without architecture modification. Sufficient experiments verify significant gains on various datasets using appealing backbones, i.e., equipped with PointCMT, PointNet++ and PointMLP achieve state-of-the-art performance on two benchmarks, i.e., 94.4% and 86.7% accuracy on ModelNet40 and ScanObjectNN, respectively. Code will be made available at https://github.com/ZhanHeshen/PointCMT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- ULIP-2: Towards Scalable Multimodal Pre-Training for 3D UnderstandingLe Xue, Ning Yu, Shu Zhang, Artemis Panagopoulou 等CVPR 2024 · 被引用 90 次
- RadOcc: Learning Cross-Modality Occupancy Knowledge through Rendering Assisted DistillationHaiming Zhang, Xu Yan, Dongfeng Bai, Jiantao Gao 等AAAI 2024 · 被引用 39 次
- Invariant Training 2D-3D Joint Hard Samples for Few-Shot Point Cloud RecognitionXuanyu Yi, Jiajun Deng, Qianru Sun, Xian-Sheng Hua 等ICCV 2023 · 被引用 17 次
- Beyond First Impressions: Integrating Joint Multi-modal Cues for Comprehensive 3D RepresentationHaowei Wang, Jiji Tang, Jiayi Ji, Xiaoshuai Sun 等ACM MM 2023 · 被引用 11 次
- FIND: Few-Shot Anomaly Inspection with Normal-Only Multi-Modal DataYiting Li, Fayao Liu, Jingyi Liao, Sichao Tian 等ICCV 2025 · 被引用 5 次
它引用的顶会 Paper22
- KPConv: Flexible and Deformable Convolution for Point CloudsHugues Thomas, Charles R. Qi, Jean-Emmanuel Deschaud, Beatriz Marcotegui 等ICCV 2019 · 被引用 3,193 次
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel 等ICCV 2019 · 被引用 2,345 次
- Revisiting Point Cloud Classification: A New Benchmark Dataset and Classification Model on Real-World DataMikaela Angelina Uy, Quang-Hieu Pham, Binh-Son Hua, Duc Thanh Nguyen 等ICCV 2019 · 被引用 1,003 次
- Rethinking Network Design and Local Geometry in Point Cloud: A Simple Residual MLP FrameworkXu Ma, Can Qin, Haoxuan You, Haoxi Ran 等ICLR 2022 · 被引用 841 次
- On the Efficacy of Knowledge DistillationJang Hyun Cho, Bharath HariharanICCV 2019 · 被引用 741 次
相关 Paper
- MM-Point: Multi-View Information-Enhanced Multi-Modal Self-Supervised 3D Point Cloud UnderstandingHai-Tao Yu, Mofei SongAAAI 2024 · 被引用 18 次
- X -Trans2Cap: Cross-Modal Knowledge Transfer using Transformer for 3D Dense CaptioningZhihao Yuan, Xu Yan, Yinghong Liao, Yao Guo 等CVPR 2022 · 被引用 72 次
- See More and Know More: Zero-shot Point Cloud Segmentation via Multi-modal Visual DataYuhang Lu, Qi Jiang, Runnan Chen, Yuenan Hou 等ICCV 2023 · 被引用 30 次
- P2P: Tuning Pre-trained Image Models for Point Cloud Analysis with Point-to-Pixel PromptingZiyi Wang, Xumin Yu, Yongming Rao, Jie Zhou 等NeurIPS 2022 · 被引用 121 次
- Take-A-Photo: 3D-to-2D Generative Pre-training of Point Cloud ModelsZiyi Wang, Xumin Yu, Yongming Rao, Jie Zhou 等ICCV 2023 · 被引用 34 次
