Simulated Annealing in Early Layers Leads to Better Generalization
AmirMohammad Sarfi, Zahra Karimpour, Muawiz Chaudhary, Nasir Mohammad Khalid, Mirco Ravanelli, Sudhir P. Mudur, Eugene Belilovsky
摘要
Recently, a number of iterative learning methods have been introduced to improve generalization. These typically rely on training for longer periods of time in exchange for improved generalization. LLF (later-layer-forgetting) is a state-of-the-art method in this category. It strengthens learning in early layers by periodically re-initializing the last few layers of the network. Our principal innovation in this work is to use Simulated annealing in EArly Layers (SEAL) of the network in place of re-initialization of later layers. Essentially, later layers go through the normal gradient descent process, while the early layers go through short stints of gradient ascent followed by gradient descent. Extensive experiments on the popular Tiny-ImageNet dataset benchmark and a series of transfer learning and few-shot learning tasks show that we outperform LLF by a significant margin. We further show that, compared to normal training, LLF features, although improving on the target task, degrade the transfer learning performance across all datasets we explored. In comparison, our method outperforms LLF across the same target datasets by a large margin. We also show that the prediction depth of our method is significantly lower than that of LLF and normal training, indicating on average better prediction performance. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- What Variables Affect Out-of-Distribution Generalization in Pretrained Models?Md Yousuf Harun, Kyungbok Lee, Gianmarco J. Gallardo, Giri Krishnan 等NeurIPS 2024 · 被引用 16 次
- Sample Selection via Contrastive Fragmentation for Noisy Label RegressionChris Dongjoo Kim, Sangwoo Moon, Jihwan Moon, Dongyeon Woo 等NeurIPS 2024 · 被引用 8 次
- Updatable Balanced Index for Fast on-Device Search with Auto-Selection ModelYushuai Ji, Sheng Wang, Zhiyu Chen, Yuan Sun 等ICDE 2026
- Gradient-Guided Annealing for Domain GeneralizationAristotelis Ballas, Christos DiouCVPR 2025
- Controlling Neural Collapse Enhances Out-of-Distribution Detection and Transfer LearningMd Yousuf Harun, Jhair Gallardo, Christopher KananICML 2025
它引用的顶会 Paper8
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- How Do Vision Transformers Work?Namuk Park, Songkuk KimICLR 2022 · 被引用 653 次
- Self-Distillation Amplifies Regularization in Hilbert SpaceHossein Mobahi, Mehrdad Farajtabar, Peter L. BartlettNeurIPS 2020 · 被引用 298 次
- What Makes Instance Discrimination Good for Transfer Learning?Nanxuan Zhao, Zhirong Wu, Rynson W. H. Lau, Stephen LinICLR 2021 · 被引用 183 次
- Self-training For Few-shot Transfer Across Extreme Task DifferencesCheng Perng Phoo, Bharath HariharanICLR 2021 · 被引用 131 次
相关 Paper
- Learning to Forget for Meta-LearningSungyong Baik, Seokil Hong, Kyoung Mu LeeCVPR 2020
- Multi-layer Rehearsal Feature Augmentation for Class-Incremental LearningBowen Zheng, Da-Wei Zhou, Han-Jia Ye, De-Chuan ZhanICML 2024 · 被引用 27 次
- RepAn: Enhanced Annealing through Re-parameterizationXiang Fei, Xiawu Zheng, Yan Wang, Fei Chao 等CVPR 2024
- RIFLE: Backpropagation in Depth for Deep Transfer Learning through Re-Initializing the Fully-connected LayErXingjian Li, Haoyi Xiong, Haozhe An, Cheng-Zhong Xu 等ICML 2020 · 被引用 45 次
- Fortuitous Forgetting in Connectionist NetworksHattie Zhou, Ankit Vani, Hugo Larochelle, Aaron C. CourvilleICLR 2022 · 被引用 50 次
