LQF: Linear Quadratic Fine-Tuning
Alessandro Achille, Aditya Golatkar, Avinash Ravichandran, Marzia Polito, Stefano Soatto
摘要
Classifiers that are linear in their parameters, and trained by optimizing a convex loss function, have predictable behavior with respect to changes in the training data, initial conditions, and optimization. Such desirable properties are absent in deep neural networks (DNNs), typically trained by non-linear fine-tuning of a pre-trained model. Previous attempts to linearize DNNs have led to interesting theoretical insights, but have not impacted the practice due to the substantial performance gap compared to standard non-linear optimization. We present the first method for linearizing a pre-trained model that achieves comparable performance to non-linear fine-tuning on most of real-world image classification tasks tested, thus enjoying the interpretability of linear models without incurring punishing losses in performance. LQF consists of simple modifications to the architecture, loss function and optimization typically used for classification: Leaky-ReLU instead of ReLU, mean squared loss instead of cross-entropy, and pre-conditioning using Kronecker factorization. None of these changes in isolation is sufficient to approach the performance of non-linear fine-tuning. When used in combination, they allow us to reach comparable performance, and even superior in the low-data regime, while enjoying the simplicity, robustness and interpretability of linear-quadratic optimization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Task Arithmetic in the Tangent Space: Improved Editing of Pre-Trained ModelsGuillermo Ortiz-Jiménez, Alessandro Favero, Pascal FrossardNeurIPS 2023 · 被引用 272 次
- TRAK: Attributing Model Behavior at ScaleSung Min Park, Kristian Georgiev, Andrew Ilyas, Guillaume Leclerc 等ICML 2023 · 被引用 260 次
- Bidirectional Learning for Offline Model-based Biological Sequence DesignCan Chen, Yingxue Zhang, Xue Liu, Mark CoatesICML 2023 · 被引用 30 次
- Tangent Model Composition for Ensembling and Continual Fine-tuningTian Yu Liu, Stefano SoattoICCV 2023 · 被引用 29 次
- Mixed Differential Privacy in Computer VisionAditya Golatkar, Alessandro Achille, Yu-Xiang Wang, Aaron Roth 等CVPR 2022 · 被引用 26 次
它引用的顶会 Paper10
- Certified Data Removal from Machine Learning ModelsChuan Guo, Tom Goldstein, Awni Y. Hannun, Laurens van der MaatenICML 2020 · 被引用 633 次
- Task2Vec: Task Embedding for Meta-LearningAlessandro Achille, Michael Lam, Rahul Tewari, Avinash Ravichandran 等ICCV 2019 · 被引用 359 次
- LEEP: A New Measure to Evaluate Transferability of Learned RepresentationsCuong V. Nguyen, Tal Hassner, Matthias W. Seeger, Cédric ArchambeauICML 2020 · 被引用 279 次
- Evaluation of Neural Architectures trained with square Loss vs Cross-Entropy in Classification TasksLike Hui, Mikhail BelkinICLR 2021 · 被引用 199 次
- Rethinking the Hyperparameters for Fine-tuningHao Li, Pratik Chaudhari, Hao Yang, Michael Lam 等ICLR 2020 · 被引用 142 次
相关 Paper
- QuadEnhancer: Leveraging Quadratic Transformations to Enhance Deep Neural NetworksQian Chen, Linxin Yang, Akang Wang, Xiaodong Luo 等NeurIPS 2025 · 被引用 2 次
- Features are fate: a theory of transfer learning in high-dimensional regressionJavan Tahir, Surya Ganguli, Grant M. RotskoffICML 2025
- Tangent Transformers for Composition, Privacy and RemovalTian Yu Liu, Aditya Golatkar, Stefano SoattoICLR 2024 · 被引用 15 次
- Scaling & Shifting Your Features: A New Baseline for Efficient Model TuningDongze Lian, Daquan Zhou, Jiashi Feng, Xinchao WangNeurIPS 2022 · 被引用 415 次
- Mixed-Privacy Forgetting in Deep NetworksAditya Golatkar, Alessandro Achille, Avinash Ravichandran, Marzia Polito 等CVPR 2021
