Composite Sketch+Text Queries for Retrieving Objects with Elusive Names and Complex Interactions
Prajwal Gatti, Kshitij Parikh, Dhriti Prasanna Paul, Manish Gupta, Anand Mishra
Abstract
Non-native speakers with limited vocabulary often struggle to name specific objects despite being able to visualize them, e.g., people outside Australia searching for ‘numbats.’ Further, users may want to search for such elusive objects with difficult-to-sketch interactions, e.g., “numbat digging in the ground.” In such common but complex situations, users desire a search interface that accepts composite multimodal queries comprising hand-drawn sketches of “difficult-to-name but easy-to-draw” objects and text describing “difficult-to-sketch but easy-to-verbalize” object's attributes or interaction with the scene. This novel problem statement distinctly differs from the previously well-researched TBIR (text-based image retrieval) and SBIR (sketch-based image retrieval) problems. To study this under-explored task, we curate a dataset, CSTBIR (Composite Sketch+Text Based Image Retrieval), consisting of 2M queries and 108K natural scene images. Further, as a solution to this problem, we propose a pretrained multimodal transformer-based baseline, STNet (Sketch+Text Network), that uses a hand-drawn sketch to localize relevant objects in the natural scene image, and encodes the text and image to perform image retrieval. In addition to contrastive learning, we propose multiple training objectives that improve the performance of our model. Extensive experiments show that our proposed method outperforms several state-of-the-art retrieval methods for text-only, sketch-only, and composite query modalities. We make the dataset and code available at: https://vl2g.github.io/projects/cstbir.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 808a23f5-93b0-4bcb-9f42-35bac75b7e55Cited by top-tier papers1
Ask how each one uses itBuilds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- Effective conditioned and composed image retrieval combining CLIP-based featuresAlberto Baldrati, Marco Bertini, Tiberio Uricchio, Alberto Del BimboCVPR 2022 · 139 citations
- Sketch3T: Test-Time Training for Zero-Shot SBIRAneeshan Sain, Ayan Kumar Bhunia, Vaishnav Potlapalli, Pinaki Nath Chowdhury et al.CVPR 2022 · 55 citations
Related papers
- Comprehensive Linguistic-Visual Composition Network for Image RetrievalHaokun Wen, Xuemeng Song, Xin Yang, Yibing Zhan et al.SIGIR 2021 · 72 citations
- TVT: Three-Way Vision Transformer through Multi-Modal Hypersphere Learning for Zero-Shot Sketch-Based Image RetrievalJialin Tian, Xing Xu, Fumin Shen, Yang Yang et al.AAAI 2022 · 54 citations
- O3SLM: Open Weight, Open Data, and Open Vocabulary Sketch-Language ModelRishi Gupta, Mukilan Karuppasamy, Shyam Marjit, Aditay Tripathi et al.AAAI 2026
- Scene-Level Sketch-Based Image Retrieval with Minimal Pairwise SupervisionCe Ge, Jingyu Wang, Qi Qi, Haifeng Sun et al.AAAI 2023 · 6 citations
- Zero-Shot Everything Sketch-Based Image Retrieval, and in Explainable StyleFengyin Lin, Mingkang Li, Da Li, Timothy M. Hospedales et al.CVPR 2023
