Fine-grained Prototypical Voting with Heterogeneous Mixup for Semi-supervised 2D-3D Cross-modal Retrieval
Fan Zhang, Xian-Sheng Hua, Chong Chen, Xiao Luo
Abstract
This paper studies the problem of semi-supervised 2D-3D retrieval, which aims to align both labeled and unla-beled 2D and 3D data into the same embedding space. The problem is challenging due to the complicated heteroge-neous relationships between 2D and 3D data. Moreover, label scarcity in real-world applications hinders from gen-erating discriminative representations. In this paper, we propose a semi-supervised approach named Fine-grained Prototypcical ⊻oting with Heterogeneous Mixup (FIVE), which maps both 2D and 3D data into a common embed-ding space for cross-modal retrieval. Specifically, we gen-erate fine-grained prototypes to model intra-class variation for both 2D and 3D data. Then, considering each unlabeled sample as a query, we retrieve relevant prototypes to vote for reliable and robust pseudo-labels, which serve as guid-ance for discriminative learning under label scarcity. Fur-thermore, to bridge the semantic gap between two modali-ties, we mix cross-modal pairs with similar semantics in the embedding space and then perform similarity learning for cross-modal discrepancy reduction in a soft manner. The whole FIVE is optimized with the consideration of sharp-ness to mitigate the impact of potential label noise. Exten-sive experiments on benchmark datasets validate the supe-riority of FIVE compared with a range of baselines in differ-ent settings. On average, FIVE outperforms the second-best approach by 4.74% on 3D MNIST, 12.94% on ModelNet10, and 22.10% on ModelNet40.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Semi-supervised Knowledge Transfer Across Multi-omic Single-cell DataFan Zhang, Tianyu Liu, Zihao Chen, Xiaojiang Peng et al.NeurIPS 2024 · 7 citations
- DiffCrossGait: Trajectory-Level Alignment for 2D-3D Cross-Modal Gait Recognition via Latent DiffusionZhiyang Lu, Ming ChengICML 2026
- Fuzzy Multimodal Learning for Trusted Cross-modal RetrievalSiyuan Duan, Yuan Sun, Dezhong Peng, Zheng Liu et al.CVPR 2025
- Intra-class Distribution-guided Generative Hashing with Neighbor Refinement for Cross-modal RetrievalHao Sun, Yadong Huo, Qibing Qin, Wenfeng Zhang et al.CVPR 2026
Builds on40
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
Related papers
- DREAM: Decoupled Discriminative Learning with Bigraph-aware Alignment for Semi-supervised 2D-3D Cross-modal RetrievalFan Zhang, Changhu Wang, Zebang Cheng, Xiaojiang Peng et al.AAAI 2025 · 1 citation
- Semi-supervised Prototype Semantic Association Learning for Robust Cross-modal RetrievalJunsheng Wang, Tiantian Gong, Yan YanSIGIR 2024 · 3 citations
- Crossmodal Few-shot 3D Point Cloud Semantic SegmentationZiyu Zhao, Zhenyao Wu, Xinyi Wu, Canyu Zhang et al.ACM MM 2022 · 20 citations
- Semi-Supervised Multi-Modal Learning with Balanced Spectral DecompositionPeng Hu, Hongyuan Zhu, Xi Peng, Jie LinAAAI 2020 · 28 citations
- More Photos Are All You Need: Semi-Supervised Learning for Fine-Grained Sketch Based Image RetrievalAyan Kumar Bhunia, Pinaki Nath Chowdhury, Aneeshan Sain, Yongxin Yang et al.CVPR 2021
