Incomplete Cross-modal Retrieval with Dual-Aligned Variational Autoencoders
Mengmeng Jing, Jingjing Li, Lei Zhu, Ke Lu, Yang Yang, Zi Huang
Abstract
Learning the relationship between the multi-modal data, e.g., texts, images and videos, is a classic task in the multimedia community. Cross-modal retrieval (CMR) is a typical example where the query and the corresponding results are in different modalities. Yet, a majority of existing works investigate CMR with an ideal assumption that the training samples in every modality are sufficient and complete. In real-world applications, however, this assumption does not always hold. Mismatch is common in multi-modal datasets. There is a high chance that samples in some modalities are either missing or corrupted. As a result, incomplete CMR has become a challenging issue. In this paper, we propose a Dual-Aligned Variational Autoencoders (DAVAE) to address the incomplete CMR problem. Specifically, we propose to learn modality-invariant representations for different modalities and use the learned representations for retrieval. We train multiple autoencoders, one for each modality, to learn the latent factors among different modalities. These latent representations are further dual-aligned at the distribution level and the semantic level to alleviate the modality gaps and enhance the discriminability of representations. For missing instances, we leverage generative models to synthesize latent representations for them. Notably, we test our method with different ratios of random incompleteness.Extensive experiments on three datasets verify that our method can consistently outperform the state-of-the-arts.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 796b2639-e442-43ed-9585-e4378ba53607Cited by top-tier papers5
- Deep Correlated Prompting for Visual Recognition with Missing ModalitiesLianyu Hu, Tongkai Shi, Wei Feng, Fanhua Shang et al.NeurIPS 2024 · 37 citations
- Dual Self-Paced Cross-Modal HashingYuan Sun, Jian Dai, Zhenwen Ren, Yingke Chen et al.AAAI 2024 · 35 citations
- Paired Cross-Modal Data Augmentation for Fine-Grained Image-to-Text RetrievalHao Wang, Guosheng Lin, Steven C. H. Hoi, Chunyan MiaoACM MM 2022 · 12 citations
- Prototype-guided Cross-modal Completion and Alignment for Incomplete Text-based Person Re-identificationTiantian Gong, Guodong Du, Junsheng Wang, Yongkang Ding et al.ACM MM 2023 · 9 citations
- Causality-Aligned Semantic Recovery for Incomplete Cross-Modal RetrievalHaipeng Chen, Yu Liu, Xun Yang, Yuheng Liang et al.AAAI 2026
Related papers
- Multimodal Disentanglement Variational AutoEncoders for Zero-Shot Cross-Modal RetrievalJialin Tian, Kai Wang, Xing Xu, Zuo Cao et al.SIGIR 2022 · 19 citations
- Associative Variational Auto-Encoder with Distributed Latent Spaces and AssociatorsDae Ung Jo, Byeongju Lee, Jongwon Choi, Haanju Yoo et al.AAAI 2020 · 8 citations
- Disentangled Cross-Modal Representation Learning with Enhanced Mutual SupervisionLu Gao, Wenlan Chen, Daoyuan Wang, Fei Guo et al.NeurIPS 2025 · 5 citations
- Dual Alignment Unsupervised Domain Adaptation for Video-Text RetrievalXiaoshuai Hao, Wanqian Zhang, Dayan Wu, Fei Zhu et al.CVPR 2023
- Uncertainty-Aware Alignment Network for Cross-Domain Video-Text RetrievalXiaoshuai Hao, Wanqian ZhangNeurIPS 2023 · 26 citations
