Towards Text-Image Interleaved Retrieval
Xin Zhang, Ziqi Dai, Yongqi Li, Yanzhao Zhang, Dingkun Long, Pengjun Xie, Meishan Zhang, Jun Yu, Wenjie Li, Min Zhang
Abstract
Current multimodal information retrieval studies mainly focus on single-image inputs, which limits real-world applications involving multiple images and text-image interleaved content. In this work, we introduce the text-image interleaved retrieval (TIIR) task, where the query and document are interleaved text-image sequences, and the model is required to understand the semantics from the interleaved context for effective retrieval. We construct a TIIR benchmark based on naturally interleaved wiki-How tutorials, where a specific pipeline is designed to generate interleaved queries. To explore the task, we adapt several off-the-shelf retrievers and build a dense baseline by interleaved multimodal large language model (MLLM). We then propose a novel Matryoshka Multimodal Embedder (MME), which compresses the number of visual tokens at different granularity, to address the challenge of excessive visual tokens in MLLM-based TIIR models. Experiments demonstrate that simple adaption of existing models does not consistently yield effective results. Our MME achieves significant improvements over the baseline by substantially fewer visual tokens. We provide extensive analysis and will release the dataset and code to facilitate future research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1a626ead-5e67-4bdb-87b7-50ed286147cdCited by top-tier papers2
- MRMR: A Realistic and Expert-Level Multidisciplinary Benchmark for Reasoning-Intensive Multimodal RetrievalSiyue Zhang, Yuan Gao, Xiao Zhou, Yilun Zhao et al.ICLR 2026 · 13 citations
- Towards Mixed-Modal Retrieval for Universal Retrieval-Augmented GenerationChenghao Zhang, Guanting Dong, Xinyu Yang, Zhicheng DouSIGIR 2026
Builds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Image Retrieval on Real-life Images with Pre-trained Vision-and-Language ModelsZheyuan Liu, Cristian Rodriguez Opazo, Damien Teney, Stephen GouldICCV 2021 · 344 citations
- Grounding Language Models to Images for Multimodal Inputs and OutputsJing Yu Koh, Ruslan Salakhutdinov, Daniel FriedICML 2023 · 160 citations
- Retrieval-Augmented Multimodal Language ModelingMichihiro Yasunaga, Armen Aghajanyan, Weijia Shi, Richard James et al.ICML 2023 · 153 citations
Related papers
- ATIR: Towards Audio-Text Interleaved Contextual RetrievalTong Zhao, Chenghao Zhang, Yutao Zhu, Zhicheng DouACL 2026
- Bridging Modalities: Improving Universal Multimodal Retrieval by Multimodal Large Language ModelsXin Zhang, Yanzhao Zhang, Wen Xie, Mingxin Li et al.CVPR 2025
- From Easy to Hard: The MIR Benchmark for Progressive Interleaved Multi-Image ReasoningHang Du, Jiayang Zhang, Guoshun Nan, Wendi Deng et al.ICCV 2025 · 1 citation
- Mm-Embed: Universal Multimodal Retrieval with Multimodal LLMSSheng-Chieh Lin, Chankyu Lee, Mohammad Shoeybi, Jimmy Lin et al.ICLR 2025
- Matryoshka Multimodal ModelsMu Cai, Jianwei Yang, Jianfeng Gao, Yong Jae LeeICLR 2025
