Orthogonal Finetuning Made Scalable
Zeju Qiu, Weiyang Liu, Adrian Weller, Bernhard Schölkopf
摘要
Orthogonal finetuning (OFT) offers highly parameter-efficient adaptation while preventing catastrophic forgetting, but its high runtime and memory demands limit practical deployment. We identify the core computational bottleneck in OFT as its weight-centric implementation, which relies on costly matrix-matrix multiplications with cubic complexity. To overcome this, we propose OFTv2, an input-centric reformulation that instead uses matrix-vector multiplications (i.e., matrix-free computation), reducing the computational cost to quadratic. We further introduce the Cayley-Neumann parameterization, an efficient orthogonal parameterization that approximates the matrix inversion in the Cayley transform via a truncated Neumann series. These modifications allow OFTv2 to achieve up to 10x faster training and 3x lower GPU memory usage without compromising performance. In addition, we extend OFTv2 to support finetuning quantized foundation models and show that it outperforms the popular QLoRA in training stability, efficiency, and memory usage.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Efficient Orthogonal Fine-Tuning with Principal Subspace AdaptationFei Wu, Jia Hu, Geyong Min, Shiqiang WangICLR 2026 · 被引用 5 次
- POET-X: Memory-efficient LLM Training by Scaling Orthogonal TransformationZeju Qiu, Lixin LIU, Adrian Weller, Han Shi 等ICML 2026 · 被引用 2 次
- Orthogonal Model MergingSihan Yang, Kexuan Shi, Weiyang LiuICML 2026 · 被引用 2 次
- Layer-Centric Factors of Variation Disentanglement for Task- and Model-Agnostic GeneralizationHee-Jun Jung, Jongmin Park, Minwoo Kang, Hoyong Kim 等ICML 2026
- Orthogonal Concept Erasure for Diffusion ModelsYuhao Sun, Lingyun Yu, Hao-Xiang Xu, Fengyuan Miao 等ICML 2026
它引用的顶会 Paper37
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context LearningHaokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta 等NeurIPS 2022 · 被引用 1,483 次
- AdaptFormer: Adapting Vision Transformers for Scalable Visual RecognitionShoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang 等NeurIPS 2022 · 被引用 1,291 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
相关 Paper
- Parameter Efficient Quasi-Orthogonal Fine-Tuning via Givens RotationXinyu Ma, Xu Chu, Zhibang Yang, Yang Lin 等ICML 2024 · 被引用 19 次
- OPLoRA: Orthogonal Projection LoRA Prevents Catastrophic Forgetting During Parameter-Efficient Fine-TuningYifeng Xiong, Xiaohui XieAAAI 2026 · 被引用 6 次
- Parameter-Efficient Orthogonal Finetuning via Butterfly FactorizationWeiyang Liu, Zeju Qiu, Yao Feng, Yuliang Xiu 等ICLR 2024 · 被引用 111 次
- Efficient Adaptation of Pre-Trained Vision Transformer Underpinned by Approximately Orthogonal Fine-Tuning StrategyYiting Yang, Hao Luo, Yuan Sun, Qingsen Yan 等ICCV 2025
- An Orthogonal High-Rank Adaptation for Large Language ModelsXin Zhang, Guang-Ze Chen, Shuzhen Li, Zhulin Liu 等EMNLP 2025
