Causality-Aligned Semantic Recovery for Incomplete Cross-Modal Retrieval
Haipeng Chen, Yu Liu, Xun Yang, Yuheng Liang, Yingda Lyu
摘要
Incomplete cross-modal retrieval (ICMR) requires models to recover missing modalities and robustly align heterogeneous ones for effective retrieval. Existing methods, however, fall short in both aspects. They often rely on limited semantic cues, such as single samples or coarse category prototypes, which compromises reconstruction quality. Moreover, these approaches are vulnerable to learning spurious cross-modal correlations, thereby impairing accurate alignment and hindering retrieval performance. To address these challenges, we propose Causality-Aligned Semantic Recovery (CASR), a novel method designed to both comprehensively restore missing modalities and mitigate spurious associations between vision and language. Our CASR involves two essential components: i) the Missing Modality Imagination (MMI) module, which combines category semantic priors with relevant contextual information to achieve high-quality semantic reconstruction; ii) the Explicit Causal Alignment (ECA) module, which explicitly learns environment-invariant attention, effectively eliminating the interference of spurious correlations and improving retrieval performance. Furthermore, we extend CASR to the challenging task of Partially Aligned Cross-Modal Retrieval, where we treat unlabeled unpaired data as a form of incomplete data. By leveraging MMI and ECA modules, we are able to learn robust representations in this setting. Extensive experiments on benchmark datasets under various missing rates demonstrate that CASR achieves superior robustness and retrieval performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Deconfounded Video Moment Retrieval with Causal InterventionXun Yang, Fuli Feng, Wei Ji, Meng Wang 等SIGIR 2021 · 被引用 198 次
- Self-Supervised Learning Disentangled Group Representation as FeatureTan Wang, Zhongqi Yue, Jianqiang Huang, Qianru Sun 等NeurIPS 2021 · 被引用 78 次
- Incomplete Cross-modal Retrieval with Dual-Aligned Variational AutoencodersMengmeng Jing, Jingjing Li, Lei Zhu, Ke Lu 等ACM MM 2020 · 被引用 63 次
- Causality-Inspired Invariant Representation Learning for Text-Based Person RetrievalYu Liu, Guihe Qin, Haipeng Chen, Zhiyong Cheng 等AAAI 2024 · 被引用 44 次
相关 Paper
- Semi-supervised Prototype Semantic Association Learning for Robust Cross-modal RetrievalJunsheng Wang, Tiantian Gong, Yan YanSIGIR 2024 · 被引用 3 次
- Partially Aligned Cross-modal Retrieval via Optimal Transport-based Prototype Alignment LearningJunsheng Wang, Tiantian Gong, Yan YanACM MM 2024 · 被引用 3 次
- Seeing What You Miss: Vision-Language Pre-training with Semantic Completion LearningYatai Ji, Rongcheng Tu, Jie Jiang, Weijie Kong 等CVPR 2023
- Recovering Coherent Affective Patterns: Addressing Modality Missing in Multimodal Sentiment AnalysisHuiting Huang, Tieliang Gong, Kai He, Wen Wen 等AAAI 2026
- DiCA: Disambiguated Contrastive Alignment for Cross-Modal Retrieval with Partial LabelsChao Su, Huiming Zheng, Dezhong Peng, Xu WangAAAI 2025 · 被引用 8 次
