Fine-grained Late-interaction Multi-modal Retrieval for Retrieval Augmented Visual Question Answering
Weizhe Lin, Jinghong Chen, Jingbiao Mei, Alexandru Coca, Bill Byrne
摘要
Knowledge-based Visual Question Answering (KB-VQA) requires VQA systems to utilize knowledge from external knowledge bases to answer visually-grounded questions. Retrieval-Augmented Visual Question Answering (RA-VQA), a strong framework to tackle KB-VQA, first retrieves related documents with Dense Passage Retrieval (DPR) and then uses them to answer questions. This paper proposes Fine-grained Late-interaction Multi-modal Retrieval (FLMR) which significantly improves knowledge retrieval in RA-VQA. FLMR addresses two major limitations in RA-VQA's retriever: (1) the image representations obtained via image-to-text transforms can be incomplete and inaccurate and (2) relevance scores between queries and documents are computed with one-dimensional embeddings, which can be insensitive to finer-grained relevance. FLMR overcomes these limitations by obtaining image representations that complement those from the image-totext transforms using a vision model aligned with an existing text-based retriever through a simple alignment network. FLMR also encodes images and questions using multi-dimensional embeddings to capture finer-grained relevance between queries and documents. FLMR significantly improves the original RA-VQA retriever's PRRecall@5 by approximately 8%. Finally, we equipped RA-VQA with two state-of-the-art large multi-modal/language models to achieve ∼ 61% VQA score in the OK-VQA dataset. * Equally contributed as the first author 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Retrieval-Augmented Visual Question Answering via Built-in Autoregressive Search EnginesXinwei Long, Zhiyuan Ma, Ermo Hua, Kaiyan Zhang 等AAAI 2025 · 被引用 18 次
- Improving Multimodal Social Media Popularity Prediction via Selective Retrieval Knowledge AugmentationXovee Xu, Yifan Zhang, Fan Zhou, Jingkuan SongAAAI 2025 · 被引用 10 次
- Highlighting What Matters: Promptable Embeddings for Attribute-Focused Image RetrievalSiting Li, Xiang Gao, Simon S. DuNeurIPS 2025 · 被引用 5 次
- LEMUR: Learned Multi-Vector RetrievalElias Jääsaari, Ville Hyvönen, Teemu RoosICML 2026 · 被引用 3 次
- Synergistic Space-Vision Processing for Predicate InferenceZhenhua Lei, Zefang Han, yu qiuICML 2026
它引用的顶会 Paper29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
相关 Paper
- PreFLMR: Scaling Up Fine-Grained Late-Interaction Multi-modal RetrieversWeizhe Lin, Jingbiao Mei, Jinghong Chen, Bill ByrneACL 2024 · 被引用 8 次
- Self-Bootstrapped Visual-Language Model for Knowledge Selection and Question AnsweringDongze Hao, Qunbo Wang, Longteng Guo, Jie Jiang 等EMNLP 2024 · 被引用 4 次
- mKG-RAG: Leveraging Multimodal Knowledge Graphs in Retrieval-Augmented Generation for Knowledge-intensive VQAXu Yuan, Liangbo Ning, Qingqing Ye, Wenqi Fan 等SIGIR 2026 · 被引用 2 次
- EntRAG: Entity-Centric Retrieval-Augmented Generation for Knowledge-based Visual Question AnsweringYiheng Hu, Xiaoyang Wang, Qing Liu, Sherry Xu 等ICML 2026
- ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question AnsweringAlberto Compagnoni, Marco Morini, Sara Sarto, Federico Cocchi 等CVPR 2026 · 被引用 11 次
