Learning Semantic Relationship among Instances for Image-Text Matching
Zheren Fu, Zhendong Mao, Yan Song, Yongdong Zhang
摘要
Image-text matching, a bridge connecting image and language, is an important task, which generally learns a holistic cross-modal embedding to achieve a high-quality semantic alignment between the two modalities. However, previous studies only focus on capturing fragmentlevel relation within a sample from a particular modality, e.g., salient regions in an image or text words in a sentence, where they usually pay less attention to capturing instance-level interactions among samples and modalities, e.g., multiple images and texts. In this paper, we argue that sample relations could help learn subtle differences for hard negative instances, and thus transfer shared knowledge for infrequent samples should be promising in obtaining better holistic embeddings. Therefore, we propose a novel hierarchical relation modeling framework (HREM), which explicitly capture both fragmentand instance-level relations to learn discriminative and robust cross-modal embeddings. Extensive experiments on Flickr30K and MS-COCO show our proposed method outperforms the state-of-the-art ones by 4%-10% in terms of rSum. Our code is available at https://github.com/ CrossmodalGroup/HREM .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper40
- Cross-modal Active Complementary Learning with Self-refining CorrespondenceYang Qin, Yuan Sun, Dezhong Peng, Joey Tianyi Zhou 等NeurIPS 2023 · 被引用 49 次
- Composing Object Relations and Attributes for Image-Text MatchingKhoi Pham, Chuong Huynh, Ser-Nam Lim, Abhinav ShrivastavaCVPR 2024 · 被引用 31 次
- Catalyst for Clustering-Based Unsupervised Object Re-identification: Feature CalibrationHuafeng Li, Qingsong Hu, Zhanxuan HuAAAI 2024 · 被引用 27 次
- CL2CM: Improving Cross-Lingual Cross-Modal Retrieval via Cross-Lingual Knowledge TransferYabing Wang, Fan Wang, Jianfeng Dong, Hao LuoAAAI 2024 · 被引用 20 次
- Contextual Augmented Global Contrast for Multimodal Intent RecognitionKaili Sun, Zhiwen Xie, Mang Ye, Huyin ZhangCVPR 2024 · 被引用 19 次
它引用的顶会 Paper17
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- Visual Semantic Reasoning for Image-Text MatchingKunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li 等ICCV 2019 · 被引用 598 次
- Similarity Reasoning and Filtration for Image-Text MatchingHaiwen Diao, Ying Zhang, Lin Ma, Huchuan LuAAAI 2021 · 被引用 413 次
- Dynamic Modality Interaction Modeling for Image-Text RetrievalLeigang Qu, Meng Liu, Jianlong Wu, Zan Gao 等SIGIR 2021 · 被引用 187 次
- Negative-Aware Attention Framework for Image-Text MatchingKun Zhang, Zhendong Mao, Quan Wang, Yongdong ZhangCVPR 2022 · 被引用 185 次
相关 Paper
- Fine-grained Image-text Matching by Cross-modal Hard Aligning NetworkZhengxin Pan, Fangyu Wu, Bailing ZhangCVPR 2023
- Maximal Matching Matters: Preventing Representation Collapse for Robust Cross-Modal RetrievalHani Alomari, Anushka Sivakumar, Andrew Zhang, Chris ThomasACL 2025
- Learning Hierarchical Cross-modal Association with Intra-modal Context for Text-Image Person RetrievalYifei Deng, Chenglong Li, Futian Wang, Jin TangACM MM 2025 · 被引用 2 次
- Bridging the Modality Gap: Dimension Information Alignment and Sparse Spatial Constraint for Image-Text MatchingXiang Ma, Xuemei Li, Lexin Fang, Caiming ZhangACM MM 2024 · 被引用 4 次
- Structured Multi-modal Feature Embedding and Alignment for Image-Sentence RetrievalXuri Ge, Fuhai Chen, Joemon M. Jose, Zhilong Ji 等ACM MM 2021 · 被引用 49 次
