Visual Query Tuning: Towards Effective Usage of Intermediate Representations for Parameter and Memory Efficient Transfer Learning
Cheng-Hao Tu, Zheda Mai, Wei-Lun Chao
摘要
Intermediate features of a pre-trained model have been shown informative for making accurate predictions on downstream tasks, even if the model backbone is kept frozen. The key challenge is how to utilize these intermediate features given their gigantic amount. We propose visual query tuning (VQT), a simple yet effective approach to aggregate intermediate features of Vision Transformers. Through introducing a handful of learnable "query" tokens to each layer, VQT leverages the inner workings of Transformers to "summarize" rich intermediate features of each layer, which can then be used to train the prediction heads of downstream tasks. As VQT keeps the intermediate features intact and only learns to combine them, it enjoys memory efficiency in training, compared to many other parameterefficient fine-tuning approaches that learn to adapt features and need back-propagation through the entire backbone. This also suggests the complementary role between VQT and those approaches in transfer learning. Empirically, VQT consistently surpasses the state-of-the-art approach that utilizes intermediate features for transfer learning and outperforms full fine-tuning in many cases. Compared to parameter-efficient approaches that adapt features, VQT achieves much higher accuracy under memory constraints. Most importantly, VQT is compatible with these approaches to attain even higher accuracy, making it a simple addon to further boost transfer learning. Code is available at https://github.com/andytu28/VQT .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper27
- What does CLIP know about a red circle? Visual prompt engineering for VLMsAleksandar Shtedritski, Christian Rupprecht, Andrea VedaldiICCV 2023 · 被引用 262 次
- Sensitivity-Aware Visual Parameter-Efficient Fine-TuningHaoyu He, Jianfei Cai, Jing Zhang, Dacheng Tao 等ICCV 2023 · 被引用 97 次
- Fine-Tuning is Fine, if CalibratedZheda Mai, Arpita Chowdhury, Ping Zhang, Cheng-Hao Tu 等NeurIPS 2024 · 被引用 34 次
- Revisiting the Power of Prompt for Visual TuningYuzhu Wang, Lechao Cheng, Chaowei Fang, Dingwen Zhang 等ICML 2024 · 被引用 33 次
- Visual-Text Cross Alignment: Refining the Similarity Score in Vision-Language ModelsJinhao Li, Haopeng Li, Sarah Monazam Erfani, Lei Feng 等ICML 2024 · 被引用 30 次
它引用的顶会 Paper29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
相关 Paper
- Fine-tuning Image Transformers using Learnable MemoryMark Sandler, Andrey Zhmoginov, Max Vladymyrov, Andrew JacksonCVPR 2022 · 被引用 51 次
- Consolidator: Mergable Adapter with Group Connections for Visual AdaptationTianxiang Hao, Hui Chen, Yuchen Guo, Guiguang DingICLR 2023 · 被引用 1 次
- WST: Wavelet-Based Multi-scale Tuning for Visual Transfer LearningJia Zeng, Lan Huang, Kangping WangAAAI 2025
- E2VPT: An Effective and Efficient Approach for Visual Prompt TuningCheng Han, Qifan Wang, Yiming Cui, Zhiwen Cao 等ICCV 2023 · 被引用 108 次
- Skip Tuning: Pre-trained Vision-Language Models are Effective and Efficient Adapters ThemselvesShihan Wu, Ji Zhang, Pengpeng Zeng, Lianli Gao 等CVPR 2025
