Generative Multi-Modal Knowledge Retrieval with Large Language Models
Xinwei Long, Jiali Zeng, Fandong Meng, Zhiyuan Ma, Kaiyan Zhang, Bowen Zhou, Jie Zhou
摘要
Knowledge retrieval with multi-modal queries plays a crucial role in supporting knowledge-intensive multi-modal applications. However, existing methods face challenges in terms of their effectiveness and training efficiency, especially when it comes to training and integrating multiple retrievers to handle multi-modal queries. In this paper, we propose an innovative end-to-end generative framework for multi-modal knowledge retrieval. Our framework takes advantage of the fact that large language models (LLMs) can effectively serve as virtual knowledge bases, even when trained with limited data. We retrieve knowledge via a two-step process: 1) generating knowledge clues related to the queries, and 2) obtaining the relevant document by searching databases using the knowledge clue. In particular, we first introduce an object-aware prefix-tuning technique to guide multi-grained visual learning. Then, we align multi-grained visual features into the textual feature space of the LLM, employing the LLM to capture cross-modal interactions. Subsequently, we construct instruction data with a unified format for model training. Finally, we propose the knowledge-guided generation strategy to impose prior constraints in the decoding steps, thereby promoting the generation of distinctive knowledge clues. Through experiments conducted on three benchmarks, we demonstrate significant improvements ranging from 3.0% to 14.6% across all evaluation metrics when compared to strong baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?Runqi Qiao, Qiuna Tan, Guanting Dong, Minhui Wu 等ACL 2025 · 被引用 236 次
- Agent4Edu: Generating Learner Response Data by Generative Agents for Intelligent Education SystemsWeibo Gao, Qi Liu, Linan Yue, Fangzhou Yao 等AAAI 2025 · 被引用 40 次
- EventVAD: Training-Free Event-Aware Video Anomaly DetectionYihua Shao, Haojin He, Sijie Li, Siyu Chen 等ACM MM 2025 · 被引用 19 次
- Collaborative Cognitive Diagnosis with Disentangled Representation Learning for Learner ModelingWeibo Gao, Qi Liu, Linan Yue, Fangzhou Yao 等NeurIPS 2024 · 被引用 12 次
- IGD: Token Decisiveness Modeling via Information Gain in LLMs for Personalized RecommendationZijie Lin, Yang Zhang, Xiaoyan Zhao, Fengbin Zhu 等NeurIPS 2025 · 被引用 12 次
它引用的顶会 Paper16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
- Transformer Memory as a Differentiable Search IndexYi Tay, Vinh Tran, Mostafa Dehghani, Jianmo Ni 等NeurIPS 2022 · 被引用 506 次
- Autoregressive Search Engines: Generating Substrings as Document IdentifiersMichele Bevilacqua, Giuseppe Ottaviano, Patrick Lewis, Scott Yih 等NeurIPS 2022 · 被引用 242 次
相关 Paper
- Generative Cross-Modal Retrieval: Memorizing Images in Multimodal Language Models for Retrieval and BeyondYongqi Li, Wenjie Wang, Leigang Qu, Liqiang Nie 等ACL 2024 · 被引用 9 次
- LLM-Enhanced Action-Aware Multi-Modal Prompt Tuning for Image-Text MatchingMengxiao Tian, Xinxiao Wu, Shuo YangICCV 2025 · 被引用 3 次
- MMRAG-RFT: Two-stage Reinforcement Fine-tuning for Explainable Multi-modal Retrieval-augmented GenerationShengwei Zhao, Jingwen Yao, Sitong Wei, Linhai Xu 等AAAI 2026
- mKG-RAG: Leveraging Multimodal Knowledge Graphs in Retrieval-Augmented Generation for Knowledge-intensive VQAXu Yuan, Liangbo Ning, Qingqing Ye, Wenqi Fan 等SIGIR 2026 · 被引用 2 次
- Multimodal Reasoning with Multimodal Knowledge GraphJunlin Lee, Yequan Wang, Jing Li, Min ZhangACL 2024 · 被引用 29 次
