5%>100%: Breaking Performance Shackles of Full Fine-Tuning on Visual Recognition Tasks
Dongshuo Yin, Leiyi Hu, Bin Li, Youqun Zhang, Xue Yang
摘要
Pre-training & fine-tuning can enhance the transferring efficiency and performance in visual tasks. Recent deltatuning methods provide more options for visual classification tasks. Despite their success, existing visual deltatuning art fails to exceed the upper limit of full finetuning on challenging tasks. To find a competitive alternative to full fine-tuning, we propose the Multi-cognitive Visual Adapter (Mona) tuning, a novel adapter-based tuning method. First, we introduce multiple vision-friendly filters into the adapter to enhance its ability for processing visual signals, while previous methods mainly rely on languagefriendly linear filters. Second, we add the scaled layernorm in the adapter to regulate the distribution of input features for visual filters. To fully demonstrate the practicality and generality of Mona, we conduct experiments on representative visual tasks, including instance segmentation on COCO, semantic segmentation on ADE20K, object detection on Pascal VOC, oriented object detection on DOTA/STAR, and image classification on three common datasets. Exciting results illustrate that Mona surpasses full fine-tuning on all these tasks by tuning less than 5% params of the backbone, and is the only delta-tuning method outperforming full fine-tuning on all tasks. For example, Mona achieves 1% performance gain on the COCO compared to full fine-tuning. Comprehensive results suggest that Monatuning is more suitable for retaining and utilizing the capabilities of pre-trained models than full fine-tuning. The code is publicly available on https://github.com/Leiyi-Hu/mona.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Drones Help Drones: A Collaborative Framework for Multi-Drone Object Trajectory Prediction and BeyondZhechao Wang, Peirui Cheng, Minxing Chen, Pengju Tian 等NeurIPS 2024 · 被引用 34 次
- UniRGB-IR: A Unified Framework for Visible-Infrared Semantic Tasks via Adapter TuningMaoxun Yuan, Bo Cui, Tianyi Zhao, Jiayi Wang 等ACM MM 2025 · 被引用 21 次
- SANSA: Unleashing the Hidden Semantics in SAM2 for Few-Shot SegmentationClaudia Cuttano, Gabriele Trivigno, Giuseppe Averta, Carlo MasoneNeurIPS 2025 · 被引用 9 次
- Dual-Kernel Adapter: Expanding Spatial Horizons for Data-Constrained Medical Image AnalysisZiquan Zhu, Hanruo Zhu, Si-Yuan Lu, Xiang Li 等ICLR 2026 · 被引用 3 次
- MP-ISMoE: Mixed-Precision Interactive Side Mixture-of-Experts for Efficient Transfer LearningYutong Zhang, Zimeng Wu, Shengcai Liao, Shujiang Wu 等AAAI 2026 · 被引用 2 次
它引用的顶会 Paper24
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao 等CVPR 2022 · 被引用 2,138 次
相关 Paper
- VMT-Adapter: Parameter-Efficient Transfer Learning for Multi-Task Dense Scene UnderstandingYi Xin, Junlong Du, Qiang Wang, Zhiwen Lin 等AAAI 2024 · 被引用 94 次
- 1% VS 100%: Parameter-Efficient Low Rank Adapter for Dense PredictionsDongshuo Yin, Yiran Yang, Zhechao Wang, Hongfeng Yu 等CVPR 2023
- Parameter-efficient is not Sufficient: Exploring Parameter, Memory, and Time Efficient Adapter Tuning for Dense PredictionsDongshuo Yin, Xueting Han, Bin Li, Hao Feng 等ACM MM 2024 · 被引用 18 次
- Vision Transformer Adapter for Dense PredictionsZhe Chen, Yuchen Duan, Wenhai Wang, Junjun He 等ICLR 2023 · 被引用 204 次
- UniAdapter: Unified Parameter-Efficient Transfer Learning for Cross-modal ModelingHaoyu Lu, Yuqi Huo, Guoxing Yang, Zhiwu Lu 等ICLR 2024 · 被引用 58 次
