Fine-grained Prototypical Voting with Heterogeneous Mixup for Semi-supervised 2D-3D Cross-modal Retrieval
Fan Zhang, Xian-Sheng Hua, Chong Chen, Xiao Luo
摘要
This paper studies the problem of semi-supervised 2D-3D retrieval, which aims to align both labeled and unla-beled 2D and 3D data into the same embedding space. The problem is challenging due to the complicated heteroge-neous relationships between 2D and 3D data. Moreover, label scarcity in real-world applications hinders from gen-erating discriminative representations. In this paper, we propose a semi-supervised approach named Fine-grained Prototypcical ⊻oting with Heterogeneous Mixup (FIVE), which maps both 2D and 3D data into a common embed-ding space for cross-modal retrieval. Specifically, we gen-erate fine-grained prototypes to model intra-class variation for both 2D and 3D data. Then, considering each unlabeled sample as a query, we retrieve relevant prototypes to vote for reliable and robust pseudo-labels, which serve as guid-ance for discriminative learning under label scarcity. Fur-thermore, to bridge the semantic gap between two modali-ties, we mix cross-modal pairs with similar semantics in the embedding space and then perform similarity learning for cross-modal discrepancy reduction in a soft manner. The whole FIVE is optimized with the consideration of sharp-ness to mitigate the impact of potential label noise. Exten-sive experiments on benchmark datasets validate the supe-riority of FIVE compared with a range of baselines in differ-ent settings. On average, FIVE outperforms the second-best approach by 4.74% on 3D MNIST, 12.94% on ModelNet10, and 22.10% on ModelNet40.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Semi-supervised Knowledge Transfer Across Multi-omic Single-cell DataFan Zhang, Tianyu Liu, Zihao Chen, Xiaojiang Peng 等NeurIPS 2024 · 被引用 7 次
- DiffCrossGait: Trajectory-Level Alignment for 2D-3D Cross-Modal Gait Recognition via Latent DiffusionZhiyang Lu, Ming ChengICML 2026
- Fuzzy Multimodal Learning for Trusted Cross-modal RetrievalSiyuan Duan, Yuan Sun, Dezhong Peng, Zheng Liu 等CVPR 2025
- Intra-class Distribution-guided Generative Hashing with Neighbor Refinement for Cross-modal RetrievalHao Sun, Yadong Huo, Qibing Qin, Wenfeng Zhang 等CVPR 2026
它引用的顶会 Paper40
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna 等NeurIPS 2020 · 被引用 7,049 次
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang 等NeurIPS 2020 · 被引用 5,129 次
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong 等NeurIPS 2020 · 被引用 2,774 次
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
相关 Paper
- DREAM: Decoupled Discriminative Learning with Bigraph-aware Alignment for Semi-supervised 2D-3D Cross-modal RetrievalFan Zhang, Changhu Wang, Zebang Cheng, Xiaojiang Peng 等AAAI 2025 · 被引用 1 次
- Semi-supervised Prototype Semantic Association Learning for Robust Cross-modal RetrievalJunsheng Wang, Tiantian Gong, Yan YanSIGIR 2024 · 被引用 3 次
- Crossmodal Few-shot 3D Point Cloud Semantic SegmentationZiyu Zhao, Zhenyao Wu, Xinyi Wu, Canyu Zhang 等ACM MM 2022 · 被引用 20 次
- Semi-Supervised Multi-Modal Learning with Balanced Spectral DecompositionPeng Hu, Hongyuan Zhu, Xi Peng, Jie LinAAAI 2020 · 被引用 28 次
- More Photos Are All You Need: Semi-Supervised Learning for Fine-Grained Sketch Based Image RetrievalAyan Kumar Bhunia, Pinaki Nath Chowdhury, Aneeshan Sain, Yongxin Yang 等CVPR 2021
