VisualSparta: An Embarrassingly Simple Approach to Large-scale Text-to-Image Search with Weighted Bag-of-words
Xiaopeng Lu, Tiancheng Zhao, Kyusong Lee
Abstract
Text-to-image retrieval is an essential task in cross-modal information retrieval, i.e., retrieving relevant images from a large and unlabelled dataset given textual queries. In this paper, we propose VisualSparta, a novel (Visualtext Sparse Transformer Matching) model that shows significant improvement in terms of both accuracy and efficiency. VisualSparta is capable of outperforming previous stateof-the-art scalable methods in MSCOCO and Flickr30K. We also show that it achieves substantial retrieving speed advantages, i.e., for a 1 million image index, VisualSparta using CPU gets ∼391X speedup compared to CPU vector search and ∼5.4X speedup compared to vector search with GPU acceleration. Experiments show that this speed advantage even gets bigger for larger datasets because Visu-alSparta can be efficiently implemented as an inverted index. To the best of our knowledge, VisualSparta is the first transformer-based textto-image retrieval model that can achieve realtime searching for large-scale datasets, with significant accuracy improvement compared to previous state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ea3f8517-c725-41f3-a1d8-5120842731b4Cited by top-tier papers3
- Chatting Makes Perfect: Chat-based Image RetrievalMatan Levy, Rami Ben-Ari, Nir Darshan, Dani LischinskiNeurIPS 2023 · 41 citations
- A Differentiable Semantic Metric Approximation in Probabilistic Embedding for Cross-Modal RetrievalHao Li, Jingkuan Song, Lianli Gao, Pengpeng Zeng et al.NeurIPS 2022 · 22 citations
- A Challenging Multimodal Video Summary: Simultaneously Extracting and Generating Keyframe-Caption Pairs from VideoKeito Kudo, Haruki Nagasawa, Jun Suzuki, Nobuyuki ShimizuEMNLP 2023 · 2 citations
Builds on4
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong et al.AAAI 2020 · 966 citations
- CAMP: Cross-Modal Adaptive Message Passing for Text-Image RetrievalZihao Wang, Xihui Liu, Hongsheng Li, Lu Sheng et al.ICCV 2019 · 349 citations
- Language-Agnostic Visual-Semantic EmbeddingsJonatas Wehrmann, Maurício Armani Lopes, Douglas M. Souza, Rodrigo C. BarrosICCV 2019 · 56 citations
- 12-in-1: Multi-Task Vision and Language Representation LearningJiasen Lu, Vedanuj Goswami, Marcus Rohrbach, Devi Parikh et al.CVPR 2020
Related papers
- Thinking Fast and Slow: Efficient Text-to-Visual Retrieval With TransformersAntoine Miech, Jean-Baptiste Alayrac, Ivan Laptev, Josef Sivic et al.CVPR 2021
- Unifying Two-Stream Encoders with Transformers for Cross-Modal RetrievalYi Bin, Haoxuan Li, Yahui Xu, Xing Xu et al.ACM MM 2023 · 33 citations
- Adaptive Cross-Modal Embeddings for Image-Text AlignmentJonatas Wehrmann, Camila Kolling, Rodrigo C. BarrosAAAI 2020 · 86 citations
- ViSTA: Vision and Scene Text Aggregation for Cross-Modal RetrievalMengjun Cheng, Yipeng Sun, Longchao Wang, Xiongwei Zhu et al.CVPR 2022 · 86 citations
- Knowledge Graph Enhanced Multimodal Transformer for Image-Text RetrievalJuncheng Zheng, Meiyu Liang, Yang Yu, Yawen Li et al.ICDE 2024 · 14 citations
