GRAM: Global Reasoning for Multi-Page VQA
Tsachi Blau, Sharon Fogel, Roi Ronen, Alona Golts, Shahar Tsiper, Elad Ben-Avraham, Aviad Aberdam, Roy Ganz, Ron Litman
Abstract
The increasing use of transformer-based large language models brings forward the challenge of processing long sequences. In document visual question answering (DocVQA), leading methods focus on the single-page setting, while documents can span hundreds of pages. We present GRAM, a method that seamlessly extends pretrained single-page models to the multi-page setting, without requiring computationally-heavy pretraining. To do so, we leverage a single-page encoder for local page-level understanding, and enhance it with document-level designated layers and learnable tokens, facilitating the flow of information across pages for global reasoning. To enforce our model to utilize the newly introduced document tokens, we propose a tailored bias adaptation method. For additional computational savings during decoding, we introduce an optional compression stage using our compressiontransformer(C-Former ),reducing the encoded sequence length, thereby allowing a tradeoff between quality and latency. Extensive experiments showcase GRAM's stateof-the-art performance on the benchmarks for multi-page DocVQA, demonstrating the effectiveness of our approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 17476c91-9c44-421b-8c4d-da574b767e7cCited by top-tier papers6
- URaG: Unified Retrieval and Generation in Multimodal LLMs for Efficient Long Document UnderstandingYongxin Shi, Jiapeng Wang, Zeyu Shan, Dezhi Peng et al.AAAI 2026 · 2 citations
- ALDEN: Reinforcement Learning for Active Navigation and Evidence Gathering in Long DocumentsTianyu Yang, Terry Ruas, Yijun Tian, Jan Philip Wahle et al.ACL 2026 · 1 citation
- DocVLM: Make Your VLM an Efficient ReaderMor Shpigel Nacson, Aviad Aberdam, Roy Ganz, Elad Ben-Avraham et al.CVPR 2025
- LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and LocatingChao Deng, Jiale Yuan, Pi Bu, Peijie Wang et al.ACL 2025
- DocVXQA: Context-Aware Visual Explanations for Document Question AnsweringMohamed Ali Souibgui, Changkyu Choi, Andrey Barsky, Kangsoo Jung et al.ICML 2025
Builds on18
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 1,168 citations
- Multimodal Few-Shot Learning with Frozen Language ModelsMaria Tsimpoukelli, Jacob Menick, Serkan Cabi, S. M. Ali Eslami et al.NeurIPS 2021 · 1,020 citations
Related papers
- CREAM: Coarse-to-Fine Retrieval and Multi-modal Efficient Tuning for Document VQAJinxu Zhang, Yongqi Yu, Yu ZhangACM MM 2024 · 4 citations
- Doc-V^*: Coarse-to-Fine Interactive Visual Reasoning for Multi-Page Document VQAYuanlei Zheng, Pei Fu, Hang Li, Ziyang Wang et al.ACL 2026 · 2 citations
- SimpleDoc: Multi-Modal Document Understanding with Dual-Cue Page Retrieval and Iterative RefinementChelsi Jain, Yiran Wu, Yifan Zeng, Jiale Liu et al.EMNLP 2025 · 1 citation
- DocPrune: Efficient Document Question Answering via Background, Question, and Comprehension-aware Token PruningJoonmyung Choi, Sanghyeok Lee, Jongha Kim, Sehyung Kim et al.CVPR 2026 · 4 citations
- DocFormer: End-to-End Transformer for Document UnderstandingSrikar Appalaraju, Bhavan Jasani, Bhargava Urala Kota, Yusheng Xie et al.ICCV 2021 · 392 citations
