Visual-Semantic Matching by Exploring High-Order Attention and Distraction
Yongzhi Li, Duo Zhang, Yadong Mu
Abstract
Cross-modality semantic matching is a vital task in computer vision and has attracted increasing attention in recent years. Existing methods mainly explore object-based alignment between image objects and text words. In this work, we address this task from two previously-ignored aspects: high-order semantic information (e.g., object-predicatesubject triplet, object-attribute pair) and visual distraction (i.e., despite the high relevance to textual query, images may also contain many prominent distracting objects or visual relations). Specifically, we build scene graphs for both visual and textual modalities. Our technical contributions are two-folds: firstly, we formulate the visual-semantic matching task as an attention-driven cross-modality scene graph matching problem. Graph convolutional networks (GCNs) are used to extract high-order information from two scene graphs. A novel cross-graph attention mechanism is proposed to contextually reweigh graph elements and calculate the inter-graph similarity; Secondly, some top-ranked samples are indeed false matching due to the co-occurrence of both highly-relevant and distracting information. We devise an information-theoretic measure for estimating semantic distraction and re-ranking the initial retrieval results. Comprehensive experiments and ablation studies on two large public datasets (MS-COCO and Flickr30K) demonstrate the superiority of the proposed method and the effectiveness of both high-order attention and distraction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9a4c38e1-fdc2-4f80-9495-7e9acfea6e98Cited by top-tier papers3
- Composing Object Relations and Attributes for Image-Text MatchingKhoi Pham, Chuong Huynh, Ser-Nam Lim, Abhinav ShrivastavaCVPR 2024 · 31 citations
- Learning to Generate Language-Supervised and Open-Vocabulary Scene Graph Using Pre-Trained Visual-Semantic SpaceYong Zhang, Yingwei Pan, Ting Yao, Rui Huang et al.CVPR 2023
- LLM4SGG: Large Language Models for Weakly Supervised Scene Graph GenerationKibum Kim, Kanghoon Yoon, Jaehyeong Jeon, Yeonjun In et al.CVPR 2024
Builds on2
Related papers
- Show Your Faith: Cross-Modal Confidence-Aware Network for Image-Text MatchingHuatian Zhang, Zhendong Mao, Kun Zhang, Yongdong ZhangAAAI 2022 · 62 citations
- Conceptual and Syntactical Cross-modal Alignment with Cross-level Consistency for Image-Text MatchingPengpeng Zeng, Lianli Gao, Xinyu Lyu, Shuaiqi Jing et al.ACM MM 2021 · 37 citations
- Hypergraph-Enhanced Hashing for Unsupervised Cross-Modal Retrieval via Robust Similarity GuidanceFangming Zhong, Chenglong Chu, Zijie Zhu, Zhikui ChenACM MM 2023 · 17 citations
- Multi-Modality Cross Attention Network for Image and Sentence MatchingXi Wei, Tianzhu Zhang, Yan Li, Yongdong Zhang et al.CVPR 2020
- Context-Aware Attention Network for Image-Text RetrievalQi Zhang, Zhen Lei, Zhaoxiang Zhang, Stan Z. LiCVPR 2020
