Retrieval-Augmented Visual Question Answering via Built-in Autoregressive Search Engines
Xinwei Long, Zhiyuan Ma, Ermo Hua, Kaiyan Zhang, Biqing Qi, Bowen Zhou
摘要
Retrieval-augmented generation (RAG) has emerged to address the knowledge-intensive visual question answering (VQA) task. Current methods mainly employ separate retrieval and generation modules to acquire external knowledge and generate answers, respectively. We propose ReAuSE, an alternative to the previous RAG model for the knowledge-based VQA task, which seamlessly integrates knowledge retriever into the generative multi-modal large language model, serving as a built-in search engine. Specifically, our model functions both as a generative retriever and an accurate answer generator. It not only helps retrieve documents from the knowledge base by producing identifier for each document, but it also answers visual questions based on the retrieved documents. Furthermore, we also propose a reinforced retrieval calibration module from relevance feedback to improve retrieval performance and align with the preferences for accurate answer generation. Extensive experiments on two representative OKVQA and A-OKVQA datasets demonstrate significant improvements ranging from 2.9% to 9.6% across all evaluation metrics when compared to strong baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- EventVAD: Training-Free Event-Aware Video Anomaly DetectionYihua Shao, Haojin He, Sijie Li, Siyu Chen 等ACM MM 2025 · 被引用 19 次
- Progressive Multimodal Search and Reasoning for Knowledge-Intensive Visual Question AnsweringChangin Choi, Wonseok Lee, Jungmin Ko, Wonjong RheeACL 2026 · 被引用 2 次
- AdsQA: Towards Advertisement Video UnderstandingXinwei Long, Kai Tian, Peng Xu, Guoli Jia 等ICCV 2025 · 被引用 1 次
- TAMEing Long Contexts in Personalization: Towards Training-Free and State-Aware MLLM Personalized AssistantRongpei Hong, Jian Lang, Ting Zhong, Yong Wang 等KDD 2026
- Memory-QA: Answering Recall Questions Based on Multimodal MemoriesHongda Jiang, Xinyuan Zhang, Siddhant Garg, Rishab Arora 等EMNLP 2025
它引用的顶会 Paper28
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch 等ICML 2023 · 被引用 2,601 次
- Benchmarking Large Language Models in Retrieval-Augmented GenerationJiawei Chen, Hongyu Lin, Xianpei Han, Le SunAAAI 2024 · 被引用 531 次
相关 Paper
- mKG-RAG: Leveraging Multimodal Knowledge Graphs in Retrieval-Augmented Generation for Knowledge-intensive VQAXu Yuan, Liangbo Ning, Qingqing Ye, Wenqi Fan 等SIGIR 2026 · 被引用 2 次
- ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question AnsweringAlberto Compagnoni, Marco Morini, Sara Sarto, Federico Cocchi 等CVPR 2026 · 被引用 11 次
- OMGM: Orchestrate Multiple Granularities and Modalities for Efficient Multimodal RetrievalWei Yang, Jingjing Fu, Rui Wang, Jinyu Wang 等ACL 2025 · 被引用 11 次
- Multi-Level Information Retrieval Augmented Generation for Knowledge-based Visual Question AnsweringOmar Adjali, Olivier Ferret, Sahar Ghannay, Hervé Le BorgneEMNLP 2024 · 被引用 4 次
- Self-Bootstrapped Visual-Language Model for Knowledge Selection and Question AnsweringDongze Hao, Qunbo Wang, Longteng Guo, Jie Jiang 等EMNLP 2024 · 被引用 4 次
