Coeff-Tuning: A Graph Filter Subspace View for Tuning Attention-Based Large Models
Zichen Miao, Wei Chen, Qiang Qiu
摘要
Transformer-based large pre-trained models have shown remarkable generalization ability, and various parameterefficient fine-tuning (PEFT) methods have been proposed to customize these models on downstream tasks with minimal computational and memory budgets. Previous PEFT methods are primarily designed from a tensor-decomposition perspective that tries to effectively tune the linear transformation by finding the smallest subset of parameters to train. Our study adopts an orthogonal view by representing the attention operation as a graph convolution and formulating the multi-head attention maps as a convolutional filter subspace, with each attention map as a subspace element. In this paper, we propose to tune the large pre-trained transformers by learning a small set of combination coefficients that construct a more expressive filter subspace from the original multi-head attention maps. We show analytically and experimentally that the tuned filter subspace can effectively expand the feature space of the multi-head attention and further enhance the capacity of transformers. We further stabilize the fine-tuning with a residual parameterization of the tunable subspace coefficients, and enhance the generalization with a regularization design by directly applying dropout on the tunable coefficient during training. The tunable coefficients take a tiny number of parameters and can be combined with previous PEFT methods in a plug-and-play manner. Extensive experiments show that our approach achieves superior performances than PEFT baselines with neglectable additional parameters. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Learning to Unlearn While Retaining: Combating Gradient Conflicts in Machine UnlearningGaurav Patel, Qiang QiuICCV 2025 · 被引用 21 次
- PEANuT: Parameter-Efficient Adaptation with Weight-aware Neural TweakersYibo Zhong, Haoxiang Jiang, Lincan Li, Ryumei Nakada 等KDD 2026 · 被引用 7 次
- Sparse Fine-Tuning of Transformers for Generative TasksWei Chen, Jingxi Yu, Zichen Miao, Qiang QiuICCV 2025 · 被引用 1 次
- Mitigating Heterogeneous Token Overfitting in LLM Knowledge EditingTianci Liu, Ruirui Li, Zihan Dong, Hui Liu 等ICML 2025
它引用的顶会 Paper47
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
相关 Paper
- Large Convolutional Model Tuning via Filter SubspaceWei Chen, Zichen Miao, Qiang QiuICLR 2025
- Generalized Tensor-Based Parameter-Efficient Fine-Tuning via Lie Group TransformationsChongjie Si, Zhiyi Shi, Xuehui Wang, Yichen Xiao 等ICCV 2025
- Efficient Orthogonal Fine-Tuning with Principal Subspace AdaptationFei Wu, Jia Hu, Geyong Min, Shiqiang WangICLR 2026 · 被引用 5 次
- WST: Wavelet-Based Multi-scale Tuning for Visual Transfer LearningJia Zeng, Lan Huang, Kangping WangAAAI 2025
- Expanding Sparse Tuning for Low Memory UsageShufan Shen, Junshu Sun, Xiangyang Ji, Qingming Huang 等NeurIPS 2024 · 被引用 12 次
