End-to-end Knowledge Retrieval with Multi-modal Queries
Man Luo, Zhiyuan Fang, Tejas Gokhale, Yezhou Yang, Chitta Baral
摘要
We investigate knowledge retrieval with multimodal queries, i.e. queries containing information split across image and text inputs, a challenging task that differs from previous work on cross-modal retrieval. We curate a new dataset called ReMuQ 1 for benchmarking progress on this task. ReMuQ requires a system to retrieve knowledge from a large corpus by integrating contents from both text and image queries. We introduce a retriever model "ReViz" that can directly process input text and images to retrieve relevant knowledge in an end-to-end fashion without being dependent on intermediate modules such as object detectors or caption generators. We introduce a new pretraining task that is effective for learning knowledge retrieval with multimodal queries and also improves performance on downstream tasks. We demonstrate superior performance in retrieval on two datasets (ReMuQ and OK-VQA) under zeroshot settings as well as further improvements when finetuned on these datasets. K1: The Empire State Building is a 102-story Art Deco skyscraper in Midtown Manhattan, New York City K2: The 828 metre (2,717 ft) tall Burj Khalifa in Dubai has been the tallest building since 2010. The Burj Khalifa has been classified as megatall. K3: The tallest building in New York is One World Trade Center which rise 1,776 feet (541 m).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Retrieval-Augmented Visual Question Answering via Built-in Autoregressive Search EnginesXinwei Long, Zhiyuan Ma, Ermo Hua, Kaiyan Zhang 等AAAI 2025 · 被引用 18 次
- Progressive Multimodal Search and Reasoning for Knowledge-Intensive Visual Question AnsweringChangin Choi, Wonseok Lee, Jungmin Ko, Wonjong RheeACL 2026 · 被引用 2 次
- Bridging Modalities: Improving Universal Multimodal Retrieval by Multimodal Large Language ModelsXin Zhang, Yanzhao Zhang, Wen Xie, Mingxin Li 等CVPR 2025
- MARA: A Multimodal Adaptive Retrieval-Augmented Framework for Document Question AnsweringHui Wu, Haoquan Zhai, Yuchen Li, Hengyi Cai 等ACM MM 2025
- Incorporating Dense Knowledge Alignment into Unified Multimodal Representation ModelsYuhao Cui, Xinxing Zu, Wenhua Zhang, Zhongzhou Zhao 等CVPR 2025
它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat 等ICML 2020 · 被引用 2,937 次
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
相关 Paper
- Reveal: Retrieval-Augmented Visual-Language Pre-Training with Multi-Source Multimodal Knowledge MemoryZiniu Hu, Ahmet Iscen, Chen Sun, Zirui Wang 等CVPR 2023
- Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?Yang Chen, Hexiang Hu, Yi Luan, Haitian Sun 等EMNLP 2023 · 被引用 37 次
- PreFLMR: Scaling Up Fine-Grained Late-Interaction Multi-modal RetrieversWeizhe Lin, Jingbiao Mei, Jinghong Chen, Bill ByrneACL 2024 · 被引用 8 次
- Uni-Retrieval: A Multi-Style Retrieval Framework for STEM's EducationYanhao Jia, Xinyi Wu, Li Hao, Qinglin Zhang 等ACL 2025
- MMCoQA: Conversational Question Answering over Text, Tables, and ImagesYongqi Li, Wenjie Li, Liqiang NieACL 2022 · 被引用 52 次
