Hi-SIGIR: Hierachical Semantic-Guided Image-to-image Retrieval via Scene Graph
Yulu Wang, Pengwen Dai, Xiaojun Jia, Zhitao Zeng, Rui Li, Xiaochun Cao
Abstract
Image-to-image retrieval, a fundamental task, aims at matching similar images based on a query image. Existing methods with convolutional neural networks are usually sensitive to low-level visual features, and ignore high-level semantic relationship information. This makes retrieving complicated images with multiple objects and various relationships a significant challenge. Although some works introduce the scene graph to capture the global semantic features of the objects and their relations, they ignore the local visual representations. In addition, due to the fragility of individual modal representations, poisoning attacks in adversarial scenarios are easily achieved, hurting the robustness of the visual-guided foundation image retrieval model. To overcome these issues, we propose a novel hierarchical semantic-guided image-to-image retrieval method via scene graph, called Hi-SIGIR. Specifically, to begin with, our proposed method generates the scene graph of an image. Then, our model extracts and learns both the visual and semantic features of the nodes and relations within the scene graphs. Next, these features are fused to obtain local information and sent to the graph neural network to obtain global information. Using these information, the similarity between the scene graphs of several images is calculated at both the local and global levels to perform image retrieval. Finally, we introduce a surrogate that calculates relevance in a cross-modal manner to understand image content better. Experimental evaluations on several wildly-used benchmarks demonstrate the superiority of the proposed method.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 8cfa13c3-7a63-4e4a-ae50-d3c9d0141058Cited by top-tier papers2
- SCENIR: Visual Semantic Clarity through Unsupervised Scene Graph RetrievalNikolaos Chaidos, Angeliki Dimitriou, Maria Lymperaiou, Giorgos StamouICML 2025
- Multi-scale Temporal Prediction via Incremental Generation and Multi-agent CollaborationZhitao Zeng, Guojian Yuan, Junyuan Mao, Yuxuan Wang et al.NeurIPS 2025
Related papers
- Image-to-Image Retrieval by Learning Similarity between Scene GraphsSangwoong Yoon, Woo-Young Kang, Sungwook Jeon, SeongEun Lee et al.AAAI 2021 · 57 citations
- Scene-Level Sketch-Based Image Retrieval with Minimal Pairwise SupervisionCe Ge, Jingyu Wang, Qi Qi, Haifeng Sun et al.AAAI 2023 · 6 citations
- Visual-Semantic Matching by Exploring High-Order Attention and DistractionYongzhi Li, Duo Zhang, Yadong MuCVPR 2020
- Scene Graph Embeddings Using Relative Similarity SupervisionParidhi Maheshwari, Ritwick Chaudhry, Vishwa VinayAAAI 2021 · 17 citations
- Domain Separation Graph Neural Networks for Saliency Object RankingZijian Wu, Jun Lu, Jing Han, Lianfa Bai et al.CVPR 2024 · 5 citations
