Skip-It? Theoretical Conditions for Layer Skipping in Vision–Language Models
Max Hartman, Vidhata Jayaraman, Moulik Choraria, Akhil Bhimaraju, Lav Varshney
Abstract
Vision–language models achieve incredible performance across a wide range of tasks, but their large size makes inference costly. Recent work has shown that multimodal processing contains significant redundancies, making it possible to skip certain layers with minimal performance loss. Yet current pruning techniques remain ad-hoc, relying on heuristics or hyperparameter sweeps rather than principled criteria for determining when layer skipping is beneficial. In this paper, we propose a unified framework that characterizes the redundancy conditions under which pruning can enhance efficiency without sacrificing performance. Central to our approach are experimentally verifiable and interpretable notions of redundancy that can be evaluated without requiring downstream task performance as a metric. Applying this framework, we corroborate prior findings that both early and late vision tokens are redundant across models, and we validate our conditions by showing they align with actual performance degradation. Beyond these empirical results, our framework provides a theoretically grounded understanding of redundancy in VLMs and unifies many of the ideas behind modern layer-skipping techniques.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0e16ee2e-ab02-4d0d-a46a-649f64055c74Cited by top-tier papers1
Ask how each one uses itBuilds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
Related papers
- VLM-Pruner: Buffering for Spatial Sparsity in an Efficient VLM Centrifugal Token Pruning ParadigmZhenkai Wu, Xiaowen Ma, Zhenliang Ni, Dengming Zhang et al.CVPR 2026 · 6 citations
- ATP-LLaVA: Adaptive Token Pruning for Large Vision Language ModelsXubing Ye, Yukang Gan, Yixiao Ge, Xiao-Ping Zhang et al.CVPR 2025
- Collaborative Multi-Mode Pruning for Vision-Language ModelsZimeng Wu, Yunhong Wang, Donghao Wang, Jiaxin ChenCVPR 2026 · 2 citations
- FlowCut: Rethinking Redundancy via Information Flow for Efficient Vision-Language ModelsJintao Tong, Wenwei Jin, Pengda Qin, Anqi Li et al.NeurIPS 2025 · 31 citations
- VFLowOpt: A Token Pruning Framework for LMMs with Visual Information Flow-Guided OptimizationSihan Yang, Runsen Xu, Chenhang Cui, Tai Wang et al.ICCV 2025 · 1 citation
