SCENIR: Visual Semantic Clarity through Unsupervised Scene Graph Retrieval
Nikolaos Chaidos, Angeliki Dimitriou, Maria Lymperaiou, Giorgos Stamou
Abstract
Despite the dominance of convolutional and transformer-based architectures in image-toimage retrieval, these models are prone to biases arising from low-level visual features, such as color. Recognizing the lack of semantic understanding as a key limitation, we propose a novel scene graph-based retrieval framework that emphasizes semantic content over superficial image characteristics. Prior approaches to scene graph retrieval predominantly rely on supervised Graph Neural Networks (GNNs), which require ground truth graph pairs driven from image captions. However, the inconsistency of captionbased supervision stemming from variable text encodings undermine retrieval reliability. To address these, we present SCENIR, a Graph Autoencoderbased unsupervised retrieval framework, which eliminates the dependence on labeled training data. Our model demonstrates superior performance across metrics and runtime efficiency, outperforming existing vision-based, multimodal, and supervised GNN approaches. We further advocate for Graph Edit Distance (GED) as a deterministic and robust ground truth measure for scene graph similarity, replacing the inconsistent caption-based alternatives for the first time in image-to-image retrieval evaluation. Finally, we validate the generalizability of our method by applying it to unannotated datasets via automated scene graph generation, while substantially contributing in advancing state-of-the-art in counterfactual image retrieval. The source code is available at https://github.com/nickhaidos/scenir-icml2025 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2cfe44ab-5baf-4ac8-a15c-57fd9fe7f47fBuilds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
- How Do Vision Transformers Work?Namuk Park, Songkuk KimICLR 2022 · 653 citations
- InceptionNeXt: When Inception Meets ConvNeXtWeihao Yu, Pan Zhou, Shuicheng Yan, Xinchao WangCVPR 2024 · 326 citations
Related papers
- Image-to-Image Retrieval by Learning Similarity between Scene GraphsSangwoong Yoon, Woo-Young Kang, Sungwook Jeon, SeongEun Lee et al.AAAI 2021 · 57 citations
- Hi-SIGIR: Hierachical Semantic-Guided Image-to-image Retrieval via Scene GraphYulu Wang, Pengwen Dai, Xiaojun Jia, Zhitao Zeng et al.ACM MM 2023 · 4 citations
- Structure Your Data: Towards Semantic Graph CounterfactualsAngeliki Dimitriou, Maria Lymperaiou, Giorgos Filandrianos, Konstantinos Thomas et al.ICML 2024 · 7 citations
- Unpaired Image Captioning via Scene Graph AlignmentsJiuxiang Gu, Shafiq R. Joty, Jianfei Cai, Handong Zhao et al.ICCV 2019 · 191 citations
- Scene-Level Sketch-Based Image Retrieval with Minimal Pairwise SupervisionCe Ge, Jingyu Wang, Qi Qi, Haifeng Sun et al.AAAI 2023 · 6 citations
