MGRAG: Semantic Subgraph Matching and Graph-Aware Caching for Multimodal Retrieval-Augmented Generation
Yubo Wang, Haoyang Li, Lei Chen
Abstract
Answering complex queries over large and heterogeneous multi-modal document corpora is a central challenge in data management, requiring fine-grained, entity-level evidence retrieval and efficient context serving. Graph-based Retrieval-Augmented Generation (RAG) systems achieve promising effectiveness by organizing multimodal documents as knowledge graphs (KGs); however, they still face three limitations: (1) query-agnostic KG construction, where corpus-wide graphs overwhelm query-relevant entities with irrelevant noise; (2) inflexible graph matching, which relies on rigid topological matching and misses path-level semantic equivalences; (3) event-agnostic KV re-computation, which scores tokens independently of graph topology, failing to preserve event-level semantic structure. To address these issues, we propose MGRAG. First, MGRAG incrementally builds query-specific KGs on demand via a lazy, top-down construction strategy. Second, we formulate graph retrieval as a path-based semantic subgraph matching problem, prove it NP-hard, and design an efficient greedy algorithm for flexible, semantics-aware retrieval. Third, MGRAG employs an event-aware KV caching mechanism to selectively recompute tokens critical to query-related events. Experiments on seven real-world multimodal QA datasets show that MGRAG achieves superior effectiveness and efficiency compared to state-of-the-art RAG, subgraph matching, and KV caching baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun et al.NeurIPS 2024 · 1,586 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Chameleon: Plug-and-Play Compositional Reasoning with Large Language ModelsPan Lu, Baolin Peng, Hao Cheng, Michel Galley et al.NeurIPS 2023 · 515 citations
Related papers
- mKG-RAG: Leveraging Multimodal Knowledge Graphs in Retrieval-Augmented Generation for Knowledge-intensive VQAXu Yuan, Liangbo Ning, Qingqing Ye, Wenqi Fan et al.SIGIR 2026 · 2 citations
- Cross-Granularity Hypergraph Retrieval-Augmented Generation for Multi-hop Question AnsweringChangjian Wang, Weihong Deng, Weili Guan, Quan Lu et al.AAAI 2026 · 2 citations
- OMGM: Orchestrate Multiple Granularities and Modalities for Efficient Multimodal RetrievalWei Yang, Jingjing Fu, Rui Wang, Jinyu Wang et al.ACL 2025 · 11 citations
- HiKEY: Hierarchical Multimodal Retrieval for Open-Domain Document Question AnsweringJoongmin Shin, Gyuho Shim, Jeongbae Park, Jaehyung Seo et al.ACL 2026
- QA-GraphRAG: Query-Adaptive Plug-and-Play Retrieval Integration for Graph-based Retrieval-Augmented GenerationZeang Sheng, Ruihong Sun, Jiahao Xu, Hanmei Luo et al.VLDB 2026
