Canonical Rank Adaptation: An Efficient Fine-Tuning Strategy for Vision Transformers
Lokesh Veeramacheneni, Moritz Wolter, Hilde Kuehne, Juergen Gall
摘要
Modern methods for fine-tuning Vision Transformers, such as Low-Rank Adaptation (LoRA) and its variants, demonstrate impressive performance. However, these methods ignore the high-dimensional nature of Multi-Head Attention (MHA) weight tensors. To address this limitation, we propose Canonical Rank Adaptation (CaRA). CaRA leverages tensor mathematics, first by tensorising the transformer into two different tensors: one for projection layers in MHA and the other for feed-forward layers. Second, the tensorised formulation is fine-tuned using the lowrank adaptation in the Canonical-Polyadic Decomposition (CPD) form. Employing CaRA efficiently minimises the number of trainable parameters. Experimentally, CaRA outperforms existing Parameter-Efficient Fine-Tuning (PEFT) methods in visual classification benchmarks such as the Visual Task Adaptation Benchmark (VTAB)-1k and the Fine-Grained Visual Categorization (FGVC) benchmark.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Beyond Static Allocation: Dynamic Sensitivity-Aware Fine-Tuning for Vision TransformersYuanyang Cao, Xichun Liu, Fuwei Zhang, Shangqi Deng 等ICML 2026
- Unpacking the Implicit Norm Dynamics of Sharpness-Aware Minimization in Tensorized ModelsTianxiao Cao, Kyohei Atarashi, Hisashi KashimaAAAI 2026
它引用的顶会 Paper10
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- Do Vision Transformers See Like Convolutional Neural Networks?Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang 等NeurIPS 2021 · 被引用 1,553 次
- AdaptFormer: Adapting Vision Transformers for Scalable Visual RecognitionShoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang 等NeurIPS 2022 · 被引用 1,291 次
相关 Paper
- Correlated Low-Rank Adaptation for ConvNetsWu Ran, Weijia Zhang, Shuyang Pang, Qi Zhu 等NeurIPS 2025 · 被引用 5 次
- TeRA: Vector-based Random Tensor Network for High-Rank Adaptation of Large Language ModelsYuxuan Gu, Wuyang Zhou, Giorgos Iacovides, Danilo P. MandicACL 2026 · 被引用 2 次
- Efficient Adaptation of Pre-trained Vision Transformer via Householder TransformationWei Dong, Yuan Sun, Yiting Yang, Xing Zhang 等NeurIPS 2024 · 被引用 10 次
- Efficient Adaptation of Pre-Trained Vision Transformer Underpinned by Approximately Orthogonal Fine-Tuning StrategyYiting Yang, Hao Luo, Yuan Sun, Qingsen Yan 等ICCV 2025
- PiCa: Parameter-Efficient Fine-Tuning with Column Space ProjectionJunseo Hwang, Wonguk Cho, Taesup KimICLR 2026 · 被引用 1 次
