VisRet: Visualization Improves Knowledge-Intensive Text-to-Image Retrieval
Di Wu, Yixin Wan, Kai-Wei Chang
Abstract
Text-to-image retrieval (T2I retrieval) remains challenging because cross-modal embeddings often behave as bags of concepts, underrepresenting structured visual relationships such as pose and viewpoint. We propose Visualize-then-Retrieve (VisRet), a retrieval paradigm that mitigates this limitation of cross-modal similarity alignment. VisRet first projects textual queries into the image modality via T2I generation, then performs retrieval within the image modality to bypass the weaknesses of cross-modal retrievers in recognizing subtle visual-spatial features. Across four benchmarks (Visual-RAG, INQUIRE-Rerank, Microsoft COCO, and our new Visual-RAG-ME featuring multientity comparisons), VisRet substantially outperforms cross-modal similarity matching and baselines that recast T2I retrieval as text-to-text similarity matching, improving nDCG@30 by 0.125 on average with CLIP as the retriever and by 0.121 with E5-V. For downstream question answering, VisRet increases accuracy on Visual-RAG and Visual-RAG-ME by 3.8% and 15.7% in top-1 retrieval, and by 3.9% and 11.1% in top-10 retrieval. Ablation studies show compatibility with different T2I instruction LLMs, T2I generation models, and downstream LLMs. VisRet provides a simple yet effective perspective for advancing in text-image retrieval. Our code and the new benchmark are publicly available at https://github.com/ xiaowu0162/Visualize-then-Retrieve . Image Embedding Image Embeddings Image Embeddings LVLM Reader Entity: Barnacle Goose Feature: unfolded wing Angle: underside visible "Generate a natural image of the underside of Barnacle Goose wings unfolded. " Visualization Instruction "When wings of Barnacle Goose (scientific name: Branta leucopsis) are folded, it displays mottled pattern. Does the same pattern appear on the underside?" Question ...
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- MuRAG: Multimodal Retrieval-Augmented Generator for Open Question Answering over Images and TextWenhu Chen, Hexiang Hu, Xi Chen, Pat Verga et al.EMNLP 2022 · 89 citations
Related papers
- OMGM: Orchestrate Multiple Granularities and Modalities for Efficient Multimodal RetrievalWei Yang, Jingjing Fu, Rui Wang, Jinyu Wang et al.ACL 2025 · 11 citations
- VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality DocumentsShi Yu, Chaoyue Tang, Bokai Xu, Junbo Cui et al.ICLR 2025
- Cross-Modal Implicit Relation Reasoning and Aligning for Text-to-Image Person RetrievalDing Jiang, Mang YeCVPR 2023
- VisualSparta: An Embarrassingly Simple Approach to Large-scale Text-to-Image Search with Weighted Bag-of-wordsXiaopeng Lu, Tiancheng Zhao, Kyusong LeeACL 2021
- Seeing Through Words: Controlling Visual Retrieval Quality with Language ModelsJianglin Lu, Simon Jenni, Kushal Kafle, Jing Shi et al.ICLR 2026 · 3 citations
