Double-Filter: Efficient Fine-tuning of Pre-trained Vision-Language Models via Patch&Layer Filtering
Yaoqin He, Junchen Fu, Kaiwen Zheng, Songpei Xu, Fuhai Chen, Jie Li, Joemon M. Jose, Xuri Ge
Abstract
In this paper, we present a novel approach, termed Double-Filter, to "slim down" the fine-tuning process of vision-language pre-trained (VLP) models via filtering redundancies in feature inputs and architectural components. We enhance the fine-tuning process using two approaches. First, we develop a new patch selection method incorporating image patch filtering through background and foreground separation, followed by a refined patch selection process. Second, we design a genetic algorithm to eliminate redundant finegrained architecture layers, improving the efficiency and effectiveness of the model. The former makes patch selection semantics more comprehensive, improving inference efficiency while ensuring semantic representation. The latter's fine-grained layer filter removes architectural redundancy to the extent possible and mitigates the impact on performance. Experimental results demonstrate that the proposed Double-Filter achieves superior efficiency of model fine-tuning and maintains competitive performance compared with the advanced efficient fine-tuning methods on three downstream tasks, VQA, NLVR and Retrieval. In addition, it has been proven to be effective under METER and ViLT VLP models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a2977a01-cae4-4a14-88c3-e18ca1f2e5d4Cited by top-tier papers2
- StructXLIP: Enhancing Vision-language Models with Multimodal Structural CuesZanxi Ruan, Songqun Gao, Qiuyu Kong, Yiming Wang et al.CVPR 2026 · 1 citation
- Vision-language Incremental Learning with Dual Class-individual MemoryFuhai Chen, Feng Zhang, Xiaoguang Ma, Yiyi Zhou et al.AAAI 2026
Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
- Towards a Unified View of Parameter-Efficient Transfer LearningJunxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick et al.ICLR 2022 · 1,182 citations
Related papers
- Parameter and Computation Efficient Transfer Learning for Vision-Language Pre-trained ModelsQiong Wu, Wei Yu, Yiyi Zhou, Shubin Huang et al.NeurIPS 2023 · 16 citations
- TRIPS: Efficient Vision-and-Language Pre-training with Text-Relevant Image Patch SelectionChaoya Jiang, Haiyang Xu, Chenliang Li, Ming Yan et al.EMNLP 2022 · 6 citations
- Too Large; Data Reduction for Vision-Language Pre-TrainingAlex Jinpeng Wang, Kevin Qinghong Lin, David Junhao Zhang, Stan Weixian Lei et al.ICCV 2023 · 35 citations
- BUS : Efficient and Effective Vision-language Pre-training with Bottom-Up Patch SummarizationChaoya Jiang, Haiyang Xu, Wei Ye, Qinghao Ye et al.ICCV 2023 · 9 citations
- MixPHM: Redundancy-Aware Parameter-Efficient Tuning for Low-Resource Visual Question AnsweringJingjing Jiang, Nanning ZhengCVPR 2023
