Canonical Rank Adaptation: An Efficient Fine-Tuning Strategy for Vision Transformers
Lokesh Veeramacheneni, Moritz Wolter, Hilde Kuehne, Juergen Gall
Abstract
Modern methods for fine-tuning Vision Transformers, such as Low-Rank Adaptation (LoRA) and its variants, demonstrate impressive performance. However, these methods ignore the high-dimensional nature of Multi-Head Attention (MHA) weight tensors. To address this limitation, we propose Canonical Rank Adaptation (CaRA). CaRA leverages tensor mathematics, first by tensorising the transformer into two different tensors: one for projection layers in MHA and the other for feed-forward layers. Second, the tensorised formulation is fine-tuned using the lowrank adaptation in the Canonical-Polyadic Decomposition (CPD) form. Employing CaRA efficiently minimises the number of trainable parameters. Experimentally, CaRA outperforms existing Parameter-Efficient Fine-Tuning (PEFT) methods in visual classification benchmarks such as the Visual Task Adaptation Benchmark (VTAB)-1k and the Fine-Grained Visual Categorization (FGVC) benchmark.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Beyond Static Allocation: Dynamic Sensitivity-Aware Fine-Tuning for Vision TransformersYuanyang Cao, Xichun Liu, Fuwei Zhang, Shangqi Deng et al.ICML 2026
- Unpacking the Implicit Norm Dynamics of Sharpness-Aware Minimization in Tensorized ModelsTianxiao Cao, Kyohei Atarashi, Hisashi KashimaAAAI 2026
Builds on10
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Do Vision Transformers See Like Convolutional Neural Networks?Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang et al.NeurIPS 2021 · 1,553 citations
- AdaptFormer: Adapting Vision Transformers for Scalable Visual RecognitionShoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang et al.NeurIPS 2022 · 1,291 citations
Related papers
- Correlated Low-Rank Adaptation for ConvNetsWu Ran, Weijia Zhang, Shuyang Pang, Qi Zhu et al.NeurIPS 2025 · 5 citations
- TeRA: Vector-based Random Tensor Network for High-Rank Adaptation of Large Language ModelsYuxuan Gu, Wuyang Zhou, Giorgos Iacovides, Danilo P. MandicACL 2026 · 2 citations
- Efficient Adaptation of Pre-trained Vision Transformer via Householder TransformationWei Dong, Yuan Sun, Yiting Yang, Xing Zhang et al.NeurIPS 2024 · 10 citations
- Efficient Adaptation of Pre-Trained Vision Transformer Underpinned by Approximately Orthogonal Fine-Tuning StrategyYiting Yang, Hao Luo, Yuan Sun, Qingsen Yan et al.ICCV 2025
- PiCa: Parameter-Efficient Fine-Tuning with Column Space ProjectionJunseo Hwang, Wonguk Cho, Taesup KimICLR 2026 · 1 citation
