DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs
Lingchen Meng, Jianwei Yang, Rui Tian, Xiyang Dai, Zuxuan Wu, Jianfeng Gao, Yu-Gang Jiang
Abstract
Most large multimodal models (LMMs) are implemented by feeding visual tokens as a sequence into the first layer of a large language model (LLM). The resulting architecture is simple but significantly increases computation and memory costs, as it has to handle a large number of additional tokens in its input layer. This paper presents a new architecture DeepStack for LMMs. Considering layers in the language and vision transformer of LMMs, we stack the visual tokens into groups and feed each group to its aligned transformer layer from bottom to top. Surprisingly, this simple method greatly enhances the power of LMMs to model interactions among visual tokens across layers but with minimal additional cost. We apply DeepStack to both language and vision transformer in LMMs, and validate the effectiveness of DeepStack LMMs with extensive empirical results. Using the same context length, our DeepStack 7B and 13B parameters surpass their counterparts by 2.7 and 2.9 on average across 9 benchmarks, respectively. Using only one-fifth of the context length, DeepStack rivals closely to the counterparts that use the full context length. These gains are particularly pronounced on high-resolution tasks, e.g., 4.2, 11.0, and 4.0 improvements on TextVQA, DocVQA, and InfoVQA compared to LLaVA-1.5-7B, respectively. We further apply DeepStack to vision transformer layers, which brings us a similar amount of improvements, 3.8 on average compared with LLaVA-1.5-7B.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1057e2f2-7720-4aff-bf69-a5ef2cf0e062Cited by top-tier papers6
- INST-IT: Boosting Instance Understanding via Explicit Visual Prompt Instruction TuningWujian Peng, Lingchen Meng, Yitong Chen, Yiweng Xie et al.NeurIPS 2025 · 7 citations
- Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene UnderstandingWencan Huang, Daizong Liu, Wei HuACM MM 2025 · 1 citation
- Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language ModelsMingyu Fu, Wei Suo, Ji Ma, Lin Yuanbo Wu et al.ACM MM 2025 · 1 citation
- Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video UnderstandingHoulun Chen, Xin Wang, Guangyao Li, Yuwei Zhou et al.SIGIR 2026
- Multi-Layer Visual Feature Fusion in Multimodal LLMs: Methods, Analysis, and Best PracticesJunyan Lin, Haoran Chen, Yue Fan, Yingqi Fan et al.CVPR 2025
Builds on24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- Multimodal Language Models See Better When They Look ShallowerHaoran Chen, Junyan Lin, Xinghao Chen, Yue Fan et al.EMNLP 2025
- Towards Interpreting Visual Information Processing in Vision-Language ModelsClement Neo, Luke Ong, Philip Torr, Mor Geva et al.ICLR 2025
- Cross-modal Information Flow in Multimodal Large Language ModelsZhi Zhang, Srishti Yadav, Fengze Han, Ekaterina ShutovaCVPR 2025
- Task-Related Token Compression in Multimodal Large Language Models from an Explainability PerspectiveLei Lei, Jie Gu, Xiaokang Ma, Chu Tang et al.ICLR 2026 · 3 citations
- Deep Pre-Alignment for VLMsTianyu Yu, Kechen Fang, Zihao Wan, Kaidong Zhang et al.ICML 2026
