SV-RAG: LoRA-Contextualizing Adaptation of MLLMs for Long Document Understanding
Jian Chen, Ruiyi Zhang, Yufan Zhou, Tong Yu, Franck Dernoncourt, Jiuxiang Gu, Ryan A. Rossi, Changyou Chen, Tong Sun
Abstract
Multimodal large language models (MLLMs) have recently shown great progress in text-rich image understanding, yet they still struggle with complex, multi-page visually-rich documents. Traditional methods using document parsers for retrievalaugmented generation suffer from performance and efficiency limitations, while directly presenting all pages to MLLMs leads to inefficiencies, especially with lengthy ones. In this work, we present a novel framework named Self-Visual Retrieval-Augmented Generation (SV-RAG ), which can broaden horizons of any MLLM to support long-document understanding. We demonstrate that MLLMs themselves can be an effective multimodal retriever to fetch relevant pages and then answer user questions based on these pages. SV-RAG is implemented with two specific MLLM adapters, one for evidence page retrieval and the other for question answering. Empirical results show state-of-the-art performance on public benchmarks, demonstrating the effectiveness of SV-RAG. * Equal contribution, work done when JC is at Adobe Research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- Resolving Evidence Sparsity: Agentic Context Engineering for Long-Document UnderstandingKeliang Liu, Zizhi Chen, Mingcheng Li, Jingqun Tang et al.CVPR 2026 · 19 citations
- DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document UnderstandingHao Yan, Yuliang Liu, Xingchen Liu, Yuyi Zhang et al.CVPR 2026 · 9 citations
- Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document UnderstandingSensen Gao, Shanshan Zhao, Xu Jiang, Lunhao Duan et al.ACL 2026 · 7 citations
- DocPrune: Efficient Document Question Answering via Background, Question, and Comprehension-aware Token PruningJoonmyung Choi, Sanghyeok Lee, Jongha Kim, Sehyung Kim et al.CVPR 2026 · 4 citations
- Multimodal LLMs as Customized Reward Models for Text-to-Image GenerationShijie Zhou, Ruiyi Zhang, Huaisheng Zhu, Branislav Kveton et al.ICCV 2025 · 3 citations
Builds on28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
Related papers
- VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality DocumentsShi Yu, Chaoyue Tang, Bokai Xu, Junbo Cui et al.ICLR 2025
- LAD-RAG: Layout-aware Dynamic RAG for Visually-Rich Document UnderstandingZhivar Sourati, Zheng Wang, Marianne Menglin Liu, Yazhe Hu et al.ACL 2026 · 5 citations
- URaG: Unified Retrieval and Generation in Multimodal LLMs for Efficient Long Document UnderstandingYongxin Shi, Jiapeng Wang, Zeyu Shan, Dezhi Peng et al.AAAI 2026 · 2 citations
- DREAM: Integrating Hierarchical Multimodal Retrieval with Multi-page Multimodal Language Model for Documents VQAJinxu Zhang, Qiyuan Fan, Yongqi Yu, Yu ZhangACM MM 2025
- MISSRAG: Addressing the Missing Modality Challenge in Multimodal Large Language ModelsVittorio Pipoli, Alessia Saporita, Federico Bolelli, Marcella Cornia et al.ICCV 2025 · 4 citations
