End-to-end Knowledge Retrieval with Multi-modal Queries
Man Luo, Zhiyuan Fang, Tejas Gokhale, Yezhou Yang, Chitta Baral
Abstract
We investigate knowledge retrieval with multimodal queries, i.e. queries containing information split across image and text inputs, a challenging task that differs from previous work on cross-modal retrieval. We curate a new dataset called ReMuQ 1 for benchmarking progress on this task. ReMuQ requires a system to retrieve knowledge from a large corpus by integrating contents from both text and image queries. We introduce a retriever model "ReViz" that can directly process input text and images to retrieve relevant knowledge in an end-to-end fashion without being dependent on intermediate modules such as object detectors or caption generators. We introduce a new pretraining task that is effective for learning knowledge retrieval with multimodal queries and also improves performance on downstream tasks. We demonstrate superior performance in retrieval on two datasets (ReMuQ and OK-VQA) under zeroshot settings as well as further improvements when finetuned on these datasets. K1: The Empire State Building is a 102-story Art Deco skyscraper in Midtown Manhattan, New York City K2: The 828 metre (2,717 ft) tall Burj Khalifa in Dubai has been the tallest building since 2010. The Burj Khalifa has been classified as megatall. K3: The tallest building in New York is One World Trade Center which rise 1,776 feet (541 m).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9cfa2565-d7ce-47a2-8036-1d15c9973763Cited by top-tier papers12
- Retrieval-Augmented Visual Question Answering via Built-in Autoregressive Search EnginesXinwei Long, Zhiyuan Ma, Ermo Hua, Kaiyan Zhang et al.AAAI 2025 · 18 citations
- Progressive Multimodal Search and Reasoning for Knowledge-Intensive Visual Question AnsweringChangin Choi, Wonseok Lee, Jungmin Ko, Wonjong RheeACL 2026 · 2 citations
- Bridging Modalities: Improving Universal Multimodal Retrieval by Multimodal Large Language ModelsXin Zhang, Yanzhao Zhang, Wen Xie, Mingxin Li et al.CVPR 2025
- MARA: A Multimodal Adaptive Retrieval-Augmented Framework for Document Question AnsweringHui Wu, Haoquan Zhai, Yuchen Li, Hengyi Cai et al.ACM MM 2025
- Incorporating Dense Knowledge Alignment into Unified Multimodal Representation ModelsYuhao Cui, Xinxing Zu, Wenhua Zhang, Zhongzhou Zhao et al.CVPR 2025
Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat et al.ICML 2020 · 2,937 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
Related papers
- Reveal: Retrieval-Augmented Visual-Language Pre-Training with Multi-Source Multimodal Knowledge MemoryZiniu Hu, Ahmet Iscen, Chen Sun, Zirui Wang et al.CVPR 2023
- Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?Yang Chen, Hexiang Hu, Yi Luan, Haitian Sun et al.EMNLP 2023 · 37 citations
- PreFLMR: Scaling Up Fine-Grained Late-Interaction Multi-modal RetrieversWeizhe Lin, Jingbiao Mei, Jinghong Chen, Bill ByrneACL 2024 · 8 citations
- Uni-Retrieval: A Multi-Style Retrieval Framework for STEM's EducationYanhao Jia, Xinyi Wu, Li Hao, Qinglin Zhang et al.ACL 2025
- MMCoQA: Conversational Question Answering over Text, Tables, and ImagesYongqi Li, Wenjie Li, Liqiang NieACL 2022 · 52 citations
