Conical Visual Concentration for Efficient Large Vision-Language Models
Long Xing, Qidong Huang, Xiaoyi Dong, Jiajie Lu, Pan Zhang, Yuhang Zang, Yuhang Cao, Conghui He, Jiaqi Wang, Feng Wu, Dahua Lin
Abstract
In large vision-language models (LVLMs), images serve as inputs that carry a wealth of information. As the idiom "A picture is worth a thousand words" implies, representing a single image in current LVLMs can require hundreds or even thousands of tokens, which thereby severely impacting the efficiency. Previous approaches have attempted to reduce the number of image tokens either before or within the early layers of LVLMs. However, these strategies inevitably result in the loss of crucial image information. To address this challenge, we conduct an empirical study revealing that all visual tokens are necessary for LVLMs in the shallow layers, and token redundancy progressively increases in the deeper layers. To this end, we propose Pyramid-Drop, a visual redundancy reduction strategy for LVLMs to boost their efficiency in both training and inference with neglectable performance loss. Specifically, we partition the LVLM into several stages and drop part of the image tokens at the end of each stage with a pre-defined ratio. The dropping is based on a lightweight similarity calculation with a negligible time overhead. Extensive experiments demonstrate that PyramidDrop can achieve over 40% training time reduction and 55% inference FLOPs acceleration on leading LVLMs like LLaVA-NeXT, maintaining comparable multi-modal performance. Besides, PyramidDrop can also serve as a plug-and-play strategy to accelerate inference in a free way, with better performance and lower inference cost than counterparts. Our code is available at https: //github.com/Cooperx521/PyramidDrop .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4d3a6756-689b-4f99-80d6-91ce1865529cCited by top-tier papers12
- IVC-Prune: Revealing the Implicit Visual Coordinates in LVLMs for Vision Token PruningZhichao Sun, Yidong Ma, Gang Liu, Nemo Chen et al.ICLR 2026 · 11 citations
- ApET: Approximation-Error Guided Token Compression for Efficient VLMsQiankun Ma, Ziyao Zhang, Haofei Wang, Zhen Song et al.CVPR 2026 · 11 citations
- Scaling the Long Video Understanding of Multimodal Large Language Models via Visual Memory MechanismTao Chen, Kun Zhang, Qiong Wu, Xiao Chen et al.CVPR 2026 · 8 citations
- ParallelVLM: Lossless Video-LLM Acceleration with Visual Alignment Aware Parallel Speculative DecodingQuan Kong, Yuhao Shen, Yicheng Ji, Huan Li et al.CVPR 2026 · 7 citations
- Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language ModelsJinlong Li, Liyuan Jiang, Haonan Zhang, Nicu SebeCVPR 2026 · 5 citations
Builds on24
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
Related papers
- Rethinking Visual Token Reduction in LVLMs Under Cross-Modal MisalignmentRui Xu, Yunke Wang, Yong Luo, Bo DuAAAI 2026 · 7 citations
- DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and InferenceAditya Kumar Singh, Hitesh Kandala, Pratik Prabhanjan Brahma, Zicheng Liu et al.CVPR 2026
- VisionZip: Longer is Better but Not Necessary in Vision Language ModelsSenqiao Yang, Yukang Chen, Zhuotao Tian, Chengyao Wang et al.CVPR 2025
- Variation-aware Vision Token Dropping for Faster Large Vision-Language ModelsChen junjie, Xuyang Liu, Zichen Wen, Yiyu Wang et al.CVPR 2026 · 24 citations
- iLLaVA: An Image is Worth Fewer Than 1/3 Input Tokens in Large Multimodal ModelsLianyu Hu, Liqing Gao, Fanhua Shang, Liang Wan et al.ICLR 2026 · 9 citations
