Unlocking the Power of Cross-Dimensional Semantic Dependency for Image-Text Matching
Kun Zhang, Lei Zhang, Bo Hu, Mengxiao Zhu, Zhendong Mao
摘要
Image-text matching, as a fundamental cross-modal task, bridges vision and language. The key challenge lies in accurately learning the semantic similarity of these two heterogeneous modalities. To determine the semantic similarity between visual and textual features, existing paradigm typically first maps them into a d-dimensional shared representation space, then independently aggregates all dimensional correspondences of cross-modal features to reflect it, e.g., the inner product. However, in this paper, we are motivated by an insightful finding that dimensions are not mutually independent, but there are intrinsic dependencies among dimensions to jointly represent latent semantics. Ignoring this intrinsic information probably leads to suboptimal aggregation for semantic similarity, impairing cross-modal matching learning. To solve this issue, we propose a novel cross-dimensional semantic dependency-aware model (called X-Dim), which explicitly and adaptively mines the semantic dependencies between dimensions in the shared space, enabling dimensions with joint dependencies to be enhanced and utilized. X-Dim (1) designs a generalized framework to learn dimensions' semantic dependency degrees, and (2) devises the adaptive sparse probabilistic learning to autonomously make the model capture precise dependencies. Theoretical analysis and extensive experiments demonstrate the superiority of X-Dim over state-of-the-art methods, achieving 5.9%-7.3% rSum improvements on Flickr30K and MS-COCO benchmarks.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper6
- Identification of Necessary Semantic Undertakers in the Causal View for Image-Text MatchingHuatian Zhang, Lei Zhang, Kun Zhang, Zhendong MaoAAAI 2024 · 被引用 12 次
- Bridging the Modality Gap: Dimension Information Alignment and Sparse Spatial Constraint for Image-Text MatchingXiang Ma, Xuemei Li, Lexin Fang, Caiming ZhangACM MM 2024 · 被引用 4 次
- Reliable Cross-modal Alignment via Prototype Iterative ConstructionXiang Ma, Litian Xu, Lexin Fang, Caiming Zhang 等ACM MM 2025 · 被引用 1 次
- Expanding the Scope of Negatives: Boosting Image-Text Matching with Negatives Distribution Guided LearningZhao Zhou, Weizhong Zhang, Xiangcheng Du, Yingbin Zheng 等AAAI 2025
- DH-Set: Improving Vision-Language Alignment with Diverse and Hybrid Set-Embeddings LearningKun Zhang, Jingyu Li, Zhe Li, S. Kevin ZhouCVPR 2025
相关 Paper
- Show Your Faith: Cross-Modal Confidence-Aware Network for Image-Text MatchingHuatian Zhang, Zhendong Mao, Kun Zhang, Yongdong ZhangAAAI 2022 · 被引用 62 次
- Multi-Modality Cross Attention Network for Image and Sentence MatchingXi Wei, Tianzhu Zhang, Yan Li, Yongdong Zhang 等CVPR 2020
- Conceptual and Syntactical Cross-modal Alignment with Cross-level Consistency for Image-Text MatchingPengpeng Zeng, Lianli Gao, Xinyu Lyu, Shuaiqi Jing 等ACM MM 2021 · 被引用 37 次
- Fine-grained Image-text Matching by Cross-modal Hard Aligning NetworkZhengxin Pan, Fangyu Wu, Bailing ZhangCVPR 2023
- Learning Semantic Relationship among Instances for Image-Text MatchingZheren Fu, Zhendong Mao, Yan Song, Yongdong ZhangCVPR 2023
