How to Fine-Tune Vision Models with SGD
Ananya Kumar, Ruoqi Shen, Sébastien Bubeck, Suriya Gunasekar
摘要
SGD and AdamW are the two most used optimizers for fine-tuning large neural networks in computer vision. When the two methods perform the same, SGD is preferable because it uses less memory (12 bytes/parameter with momentum and 8 bytes/parameter without) than AdamW (16 bytes/parameter). However, on a suite of downstream tasks, especially those with distribution shifts, we find that fine-tuning with AdamW performs substantially better than SGD on modern Vision Transformer and ConvNeXt models. We find that large gaps in performance between SGD and AdamW occur when the fine-tuning gradients in the first "embedding" layer are much larger than in the rest of the model. Our analysis suggests an easy fix that works consistently across datasets and models: freezing the embedding layer (less than 1% of the parameters) leads to SGD with or without momentum performing slightly better than AdamW while using less memory (e.g., on ViT-L, SGD uses ∼ 33% less GPU memory). Our insights result in state-ofthe-art accuracies on five popular distribution shift benchmarks: WILDS-FMoW, WILDS-Camelyon, BREEDS-Living-17, Waterbirds, and DomainNet.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Task-Specific Skill Localization in Fine-tuned Language ModelsAbhishek Panigrahi, Nikunj Saunshi, Haoyu Zhao, Sanjeev AroraICML 2023 · 被引用 100 次
- Surgical Fine-Tuning Improves Adaptation to Distribution ShiftsYoonho Lee, Annie S. Chen, Fahim Tajwar, Ananya Kumar 等ICLR 2023 · 被引用 47 次
- Out-of-Domain Robustness via Targeted AugmentationsIrena Gao, Shiori Sagawa, Pang Wei Koh, Tatsunori Hashimoto 等ICML 2023 · 被引用 33 次
- On the Implicit Bias of AdamMatias D. Cattaneo, Jason M. Klusowski, Boris ShigidaICML 2024 · 被引用 26 次
- Trainable Transformer in TransformerAbhishek Panigrahi, Sadhika Malladi, Mengzhou Xia, Sanjeev AroraICML 2024 · 被引用 16 次
它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
相关 Paper
- Do We Need Adam? Surprisingly Strong and Sparse Reinforcement Learning with SGD in LLMsSagnik Mukherjee, Lifan Yuan, Pavan Jayasinha, Dilek Hakkani-Tür 等ICML 2026 · 被引用 4 次
- Early Convolutions Help Transformers See BetterTete Xiao, Mannat Singh, Eric Mintun, Trevor Darrell 等NeurIPS 2021 · 被引用 974 次
- EfficientFSL: Enhancing Few-Shot Classification via Query-Only Tuning In Vision TransformersWenwen Liao, Hang Ruan, Jianbo Yu, Bing Song 等AAAI 2026
- Fine-tuning Image Transformers using Learnable MemoryMark Sandler, Andrey Zhmoginov, Max Vladymyrov, Andrew JacksonCVPR 2022 · 被引用 51 次
- ZO-AdaMU Optimizer: Adapting Perturbation by the Momentum and Uncertainty in Zeroth-Order OptimizationShuoran Jiang, Qingcai Chen, Youcheng Pan, Yang Xiang 等AAAI 2024 · 被引用 27 次
