Telling the What while Pointing to the Where: Multimodal Queries for Image Retrieval
Soravit Changpinyo, Jordi Pont-Tuset, Vittorio Ferrari, Radu Soricut
Abstract
Most existing image retrieval systems use text queries as a way for the user to express what they are looking for. However, fine-grained image retrieval often requires the ability to also express where in the image the content they are looking for is. The text modality can only cumbersomely express such localization preferences, whereas pointing is a more natural fit. In this paper, we propose an image retrieval setup with a new form of multimodal queries, where the user simultaneously uses both spoken natural language (the what) and mouse traces over an empty canvas (the where) to express the characteristics of the desired target image. We then describe simple modifications to an existing image retrieval model, enabling it to operate in this setup. Qualitative and quantitative experiments show that our model effectively takes this spatial guidance into account, and provides significantly more accurate retrieval results compared to text-only equivalent systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b36bf24a-0a0b-477b-b15a-8d5ce663fd4eCited by top-tier papers6
- Progressive Spatio-Temporal Prototype Matching for Text-Video RetrievalPandeng Li, Chen-Wei Xie, Liming Zhao, Hongtao Xie et al.ICCV 2023 · 62 citations
- Who are you referring to? Coreference resolution in image narrationsArushi Goel, Basura Fernando, Frank Keller, Hakan BilenICCV 2023 · 8 citations
- Composite Sketch+Text Queries for Retrieving Objects with Elusive Names and Complex InteractionsPrajwal Gatti, Kshitij Parikh, Dhriti Prasanna Paul, Manish Gupta et al.AAAI 2024 · 6 citations
- EDIS: Entity-Driven Image Search over Multimodal Web ContentSiqi Liu, Weixi Feng, Tsu-Jui Fu, Wenhu Chen et al.EMNLP 2023 · 6 citations
- GENIUS: A Generative Framework for Universal Multimodal SearchSungyeon Kim, Xinliang Zhu, Xiaofan Lin, Muhammet Bastan et al.CVPR 2025
Builds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong et al.AAAI 2020 · 966 citations
- Visual Semantic Reasoning for Image-Text MatchingKunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li et al.ICCV 2019 · 598 citations
- Similarity Reasoning and Filtration for Image-Text MatchingHaiwen Diao, Ying Zhang, Lin Ma, Huchuan LuAAAI 2021 · 413 citations
Related papers
- Composed Query Image Retrieval Using Locally Bounded FeaturesMehrdad Hosseinzadeh, Yang WangCVPR 2020
- Cross-modal Joint Prediction and Alignment for Composed Query Image RetrievalYuchen Yang, Min Wang, Wengang Zhou, Houqiang LiACM MM 2021 · 29 citations
- Learning GUI Grounding with Spatial Reasoning from Visual FeedbackYu Zhao, Wei-Ning Chen, Huseyin Inan, Samuel Kessler et al.ICML 2026 · 9 citations
- Comprehensive Linguistic-Visual Composition Network for Image RetrievalHaokun Wen, Xuemeng Song, Xin Yang, Yibing Zhan et al.SIGIR 2021 · 72 citations
- You'll Never Walk Alone: A Sketch and Text Duet for Fine-Grained Image RetrievalSubhadeep Koley, Ayan Kumar Bhunia, Aneeshan Sain, Pinaki Nath Chowdhury et al.CVPR 2024
