LQF: Linear Quadratic Fine-Tuning
Alessandro Achille, Aditya Golatkar, Avinash Ravichandran, Marzia Polito, Stefano Soatto
Abstract
Classifiers that are linear in their parameters, and trained by optimizing a convex loss function, have predictable behavior with respect to changes in the training data, initial conditions, and optimization. Such desirable properties are absent in deep neural networks (DNNs), typically trained by non-linear fine-tuning of a pre-trained model. Previous attempts to linearize DNNs have led to interesting theoretical insights, but have not impacted the practice due to the substantial performance gap compared to standard non-linear optimization. We present the first method for linearizing a pre-trained model that achieves comparable performance to non-linear fine-tuning on most of real-world image classification tasks tested, thus enjoying the interpretability of linear models without incurring punishing losses in performance. LQF consists of simple modifications to the architecture, loss function and optimization typically used for classification: Leaky-ReLU instead of ReLU, mean squared loss instead of cross-entropy, and pre-conditioning using Kronecker factorization. None of these changes in isolation is sufficient to approach the performance of non-linear fine-tuning. When used in combination, they allow us to reach comparable performance, and even superior in the low-data regime, while enjoying the simplicity, robustness and interpretability of linear-quadratic optimization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Task Arithmetic in the Tangent Space: Improved Editing of Pre-Trained ModelsGuillermo Ortiz-Jiménez, Alessandro Favero, Pascal FrossardNeurIPS 2023 · 272 citations
- TRAK: Attributing Model Behavior at ScaleSung Min Park, Kristian Georgiev, Andrew Ilyas, Guillaume Leclerc et al.ICML 2023 · 260 citations
- Bidirectional Learning for Offline Model-based Biological Sequence DesignCan Chen, Yingxue Zhang, Xue Liu, Mark CoatesICML 2023 · 30 citations
- Tangent Model Composition for Ensembling and Continual Fine-tuningTian Yu Liu, Stefano SoattoICCV 2023 · 29 citations
- Mixed Differential Privacy in Computer VisionAditya Golatkar, Alessandro Achille, Yu-Xiang Wang, Aaron Roth et al.CVPR 2022 · 26 citations
Builds on10
- Certified Data Removal from Machine Learning ModelsChuan Guo, Tom Goldstein, Awni Y. Hannun, Laurens van der MaatenICML 2020 · 633 citations
- Task2Vec: Task Embedding for Meta-LearningAlessandro Achille, Michael Lam, Rahul Tewari, Avinash Ravichandran et al.ICCV 2019 · 359 citations
- LEEP: A New Measure to Evaluate Transferability of Learned RepresentationsCuong V. Nguyen, Tal Hassner, Matthias W. Seeger, Cédric ArchambeauICML 2020 · 279 citations
- Evaluation of Neural Architectures trained with square Loss vs Cross-Entropy in Classification TasksLike Hui, Mikhail BelkinICLR 2021 · 199 citations
- Rethinking the Hyperparameters for Fine-tuningHao Li, Pratik Chaudhari, Hao Yang, Michael Lam et al.ICLR 2020 · 142 citations
Related papers
- QuadEnhancer: Leveraging Quadratic Transformations to Enhance Deep Neural NetworksQian Chen, Linxin Yang, Akang Wang, Xiaodong Luo et al.NeurIPS 2025 · 2 citations
- Features are fate: a theory of transfer learning in high-dimensional regressionJavan Tahir, Surya Ganguli, Grant M. RotskoffICML 2025
- Tangent Transformers for Composition, Privacy and RemovalTian Yu Liu, Aditya Golatkar, Stefano SoattoICLR 2024 · 15 citations
- Scaling & Shifting Your Features: A New Baseline for Efficient Model TuningDongze Lian, Daquan Zhou, Jiashi Feng, Xinchao WangNeurIPS 2022 · 415 citations
- Mixed-Privacy Forgetting in Deep NetworksAditya Golatkar, Alessandro Achille, Avinash Ravichandran, Marzia Polito et al.CVPR 2021
