Zero-shot Multimodal Document Retrieval via Cross-modal Question Generation
Yejin Choi, Jae-Woo Park, Janghan Yoon, Saejin Kim, Jaehyun Jeon, Youngjae Yu
Abstract
Rapid advances in Multimodal Large Language Models (MLLMs) have extended information retrieval beyond text, enabling access to complex real-world documents that combine both textual and visual content. However, most documents are private, either owned by individuals or confined within corporate silos, and current retrievers struggle when faced with unseen domains or languages. To address this gap, we introduce PREMIR, a simple yet effective framework that leverages the broad knowledge of an MLLM to generate cross-modal pre-questions (preQs) before retrieval. Unlike earlier multimodal retrievers that embed entire documents as a single vector, PREMIR leverages preQs, decomposed from documents into finer token-level representations across modalities, enabling richer contextual understanding. Experiments show that PREMIR achieves stateof-the-art performance on out-of-distribution benchmarks, including closed-domain and multilingual settings, outperforming strong baselines across all metrics. We confirm the contribution of each component through in-depth ablation studies, and qualitative analyses of the generated preQs further highlight the framework's robustness in real-world settings 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 82a2df60-8b04-4068-8486-cb6b3bf199acCited by top-tier papers1
Ask how each one uses itBuilds on13
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- Large Language Models Are Zero-Shot Time Series ForecastersNate Gruver, Marc Finzi, Shikai Qiu, Andrew Gordon WilsonNeurIPS 2023 · 898 citations
Related papers
- PreFLMR: Scaling Up Fine-Grained Late-Interaction Multi-modal RetrieversWeizhe Lin, Jingbiao Mei, Jinghong Chen, Bill ByrneACL 2024 · 8 citations
- Mm-Embed: Universal Multimodal Retrieval with Multimodal LLMSSheng-Chieh Lin, Chankyu Lee, Mohammad Shoeybi, Jimmy Lin et al.ICLR 2025
- Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-trainingJunlin Han, Shengbang Tong, David Fan, Yufan Ren et al.ICLR 2026 · 25 citations
- FreeRet: MLLMs as Training-Free RetrieversYuhan Zhu, Xiangyu Zeng, Chenting Wang, Xinhao Li et al.ICML 2026 · 5 citations
- Multimodal LLM Enhanced Cross-lingual Cross-modal RetrievalYabing Wang, Le Wang, Qiang Zhou, Zhibin Wang et al.ACM MM 2024 · 24 citations
