Conceptual and Syntactical Cross-modal Alignment with Cross-level Consistency for Image-Text Matching
Pengpeng Zeng, Lianli Gao, Xinyu Lyu, Shuaiqi Jing, Jingkuan Song
Abstract
Image-Text Matching (ITM) is a fundamental and emerging task, which plays a key role in cross-modal understanding. It remains a challenge because prior works mainly focus on learning fine-grained (i.e. coarse and/or phrase) correspondence, without considering the syntactical correspondence. In theory, a sentence is not only a set of words or phrases but also a syntactic structure, consisting of a set of basic syntactic tuples (i.e.(attribute) object - predicate - (attribute) subject). Inspired by this, we propose a Conceptual and Syntactical Cross-modal Alignment with Cross-level Consistency (CSCC) for Image-text Matching by simultaneously exploring the multiple-level cross-modal alignments across the concept and syntactic with a consistency constraint. Specifically, a conceptual-level cross-modal alignment is introduced for exploring the fine-grained correspondence, while a syntactical-level cross-modal alignment is proposed to explicitly learn a high-level syntactic similarity function. Moreover, an empirical cross-level consistent attention loss is introduced to maintain the consistency between cross-modal attentions obtained from the above two cross-modal alignments. To justify our method, comprehensive experiments are conducted on two public benchmark datasets, i.e. MS-COCO (1K and 5K) and Flickr30K, which show that our CSCC outperforms state-of-the-art methods with fairly competitive improvements.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get c478b6a6-d877-4ca0-b1cc-b5f80614a849Cited by top-tier papers8
- Prototype-based Aleatoric Uncertainty Quantification for Cross-modal RetrievalHao Li, Jingkuan Song, Lianli Gao, Xiaosu Zhu et al.NeurIPS 2023 · 52 citations
- Fine-Grained Predicates Learning for Scene Graph GenerationXinyu Lyu, Lianli Gao, Yuyu Guo, Zhou Zhao et al.CVPR 2022 · 48 citations
- Progressive Tree-Structured Prototype Network for End-to-End Image CaptioningPengpeng Zeng, Jinkuan Zhu, Jingkuan Song, Lianli GaoACM MM 2022 · 35 citations
- Cross-Lingual Cross-Modal Retrieval with Noise-Robust LearningYabing Wang, Jianfeng Dong, Tianxiang Liang, Minsong Zhang et al.ACM MM 2022 · 26 citations
- A Differentiable Semantic Metric Approximation in Probabilistic Embedding for Cross-Modal RetrievalHao Li, Jingkuan Song, Lianli Gao, Pengpeng Zeng et al.NeurIPS 2022 · 22 citations
Related papers
- Multi-Modality Cross Attention Network for Image and Sentence MatchingXi Wei, Tianzhu Zhang, Yan Li, Yongdong Zhang et al.CVPR 2020
- Graph Structured Network for Image-Text MatchingChunxiao Liu, Zhendong Mao, Tianzhu Zhang, Hongtao Xie et al.CVPR 2020
- Fine-grained Image-text Matching by Cross-modal Hard Aligning NetworkZhengxin Pan, Fangyu Wu, Bailing ZhangCVPR 2023
- Show Your Faith: Cross-Modal Confidence-Aware Network for Image-Text MatchingHuatian Zhang, Zhendong Mao, Kun Zhang, Yongdong ZhangAAAI 2022 · 62 citations
- CoV-Align: Efficient Fine-grained Cross-Modal Alignment with Cohesive Visual Semantics PriorityHengqi Liu, Wanting Zhou, Longteng Kong, Fangxiang Feng et al.CVPR 2026
