Scene-Level Sketch-Based Image Retrieval with Minimal Pairwise Supervision
Ce Ge, Jingyu Wang, Qi Qi, Haifeng Sun, Tong Xu, Jianxin Liao
Abstract
The sketch-based image retrieval (SBIR) task has long been researched at the instance level, where both query sketches and candidate images are assumed to contain only one dominant object. This strong assumption constrains its application, especially with the increasingly popular intelligent terminals and human-computer interaction technology. In this work, a more general scene-level SBIR task is explored, where sketches and images can both contain multiple object instances. The new general task is extremely challenging due to several factors: (i) scene-level SBIR inherently shares sketch-specific difficulties with instance-level SBIR (e.g., sparsity, abstractness, and diversity), (ii) the cross-modal similarity is measured between two partially aligned domains (i.e., not all objects in images are drawn in scene sketches), and (iii) besides instance-level visual similarity, a more complex multi-dimensional scene-level feature matching problem is imposed (including appearance, semantics, layout, etc.). Addressing these challenges, a novel Conditional Graph Autoencoder model is proposed to deal with scene-level sketch-images retrieval. More importantly, the model can be trained with only pairwise supervision, which distinguishes our study from others in that elaborate instance-level annotations (for example, bounding boxes) are no longer required. Extensive experiments confirm the ability of our model to robustly retrieve multiple related objects at the scene level and exhibit superior performance beyond strong competitors.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 76da4ec5-fa33-4ecc-abfd-543ccbe663edCited by top-tier papers2
- Multi-Modal Interactive Agent Layer for Few-Shot Universal Cross-Domain Retrieval and BeyondKaixiang Chen, Pengfei Fang, Hui XueNeurIPS 2025 · 4 citations
- DePro: Domain Ensemble using Decoupled Prompts for Universal Cross-Domain RetrievalKaixiang Chen, Pengfei Fang, Hui XueSIGIR 2025 · 2 citations
Builds on1
Related papers
- StyleMeUp: Towards Style-Agnostic Sketch-Based Image RetrievalAneeshan Sain, Ayan Kumar Bhunia, Yongxin Yang, Tao Xiang et al.CVPR 2021
- Zero-Shot Sketch-Based Image Retrieval via Graph Convolution NetworkZhaolong Zhang, Yuejie Zhang, Rui Feng, Tao Zhang et al.AAAI 2020 · 67 citations
- Hi-SIGIR: Hierachical Semantic-Guided Image-to-image Retrieval via Scene GraphYulu Wang, Pengwen Dai, Xiaojun Jia, Zhitao Zeng et al.ACM MM 2023 · 4 citations
- Sketch Less for More: On-the-Fly Fine-Grained Sketch-Based Image RetrievalAyan Kumar Bhunia, Yongxin Yang, Timothy M. Hospedales, Tao Xiang et al.CVPR 2020
- SCENIR: Visual Semantic Clarity through Unsupervised Scene Graph RetrievalNikolaos Chaidos, Angeliki Dimitriou, Maria Lymperaiou, Giorgos StamouICML 2025
