Semi-Supervised Multi-Modal Learning with Balanced Spectral Decomposition
Peng Hu, Hongyuan Zhu, Xi Peng, Jie Lin
摘要
Cross-modal retrieval aims to retrieve the relevant samples across different modalities, of which the key problem is how to model the correlations among different modalities while narrowing the large heterogeneous gap. In this paper, we propose a Semi-supervised Multimodal Learning Network method (SMLN) which correlates different modalities by capturing the intrinsic structure and discriminative correlation of the multimedia data. To be specific, the labeled and unlabeled data are used to construct a similarity matrix which integrates the cross-modal correlation, discrimination, and intra-modal graph information existing in the multimedia data. What is more important is that we propose a novel optimization approach to optimize our loss within a neural network which involves a spectral decomposition problem derived from a ratio trace criterion. Our optimization enjoys two advantages given below. On the one hand, the proposed approach is not limited to our loss, which could be applied to any case that is a neural network with the ratio trace criterion. On the other hand, the proposed optimization is different from existing ones which alternatively maximize the minor eigenvalues, thus overemphasizing the minor eigenvalues and ignore the dominant ones. In contrast, our method will exactly balance all eigenvalues, thus being more competitive to existing methods. Thanks to our loss and optimization strategy, our method could well preserve the discriminative and instinct information into the common space and embrace the scalability in handling large-scale multimedia data. To verify the effectiveness of the proposed method, extensive experiments are carried out on three widely-used multimodal datasets comparing with 13 state-of-the-art approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Semi-IIN: Semi-Supervised Intra-Inter Modal Interaction Learning Network for Multimodal Sentiment AnalysisJinhao Lin, Yifei Wang, Yanwu Xu, Qi LiuAAAI 2025 · 被引用 3 次
- Order-Preserving Dimension Reduction for Multimodal Semantic EmbeddingChengyu Gong, Gefei Shen, Luanzheng Guo, Nathan R. Tallent 等AAAI 2026 · 被引用 2 次
- Learning Cross-Modal Retrieval With Noisy LabelsPeng Hu, Xi Peng, Hongyuan Zhu, Liangli Zhen 等CVPR 2021
- Fuzzy Multimodal Learning for Trusted Cross-modal RetrievalSiyuan Duan, Yuan Sun, Dezhong Peng, Zheng Liu 等CVPR 2025
相关 Paper
- Multi-graph Convolutional Network for Unsupervised 3D Shape RetrievalWeizhi Nie, Yue Zhao, An-An Liu, Zan Gao 等ACM MM 2020 · 被引用 10 次
- MCCN: Multimodal Coordinated Clustering Network for Large-Scale Cross-modal RetrievalZhixiong Zeng, Ying Sun, Wenji MaoACM MM 2021 · 被引用 20 次
- C3CMR: Cross-Modality Cross-Instance Contrastive Learning for Cross-Media RetrievalJunsheng Wang, Tiantian Gong, Zhixiong Zeng, Changchang Sun 等ACM MM 2022 · 被引用 12 次
- DREAM: Decoupled Discriminative Learning with Bigraph-aware Alignment for Semi-supervised 2D-3D Cross-modal RetrievalFan Zhang, Changhu Wang, Zebang Cheng, Xiaojiang Peng 等AAAI 2025 · 被引用 1 次
- Semi-supervised Prototype Semantic Association Learning for Robust Cross-modal RetrievalJunsheng Wang, Tiantian Gong, Yan YanSIGIR 2024 · 被引用 3 次
