Trainable Weight Averaging: Efficient Training by Optimizing Historical Solutions
Tao Li, Zhehao Huang, Qinghua Tao, Yingwen Wu, Xiaolin Huang
摘要
Stochastic gradient descent (SGD) and its variants are considered as the de-facto methods to train deep neural networks (DNNs). While recent improvements to SGD mainly focus on the descent algorithm itself, few works pay attention to utilizing the historical solutions---as an iterative method, SGD has gone through substantial explorations before convergence. Recently, an interesting attempt is stochastic weight averaging (SWA), which significantly improves the generalization by simply averaging the solutions at the tail stage of training. In this paper, we realize that the averaging coefficients could be determined in a trainable manner and propose Trainable Weight Averaging (TWA), a novel optimization method in the reduced subspace spanned by historical solutions. TWA has much greater flexibility and can be applied to the head stage of training to achieve training efficiency while preserving good generalization capability. Further, we propose a distributed training scheme to resolve the memory burden of large-scale training with efficient parallel computation. In the extensive numerical experiments, (i) TWA achieves consistent improvements over SWA with less sensitivity to learning rate; (ii) applying TWA in the head stage of training largely speeds up the convergence, resulting in over time saving on CIFAR and on ImageNet with improved generalization compared with regular training.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper13
- Representation Surgery for Multi-Task Model MergingEnneng Yang, Li Shen, Zhenyi Wang, Guibing Guo 等ICML 2024 · 被引用 96 次
- Model Merging in Pre-training of Large Language ModelsYunshui Li, Yiyuan Ma, Shen Yan, Chaoyi Zhang 等NeurIPS 2025 · 被引用 40 次
- WSM: Decay-Free Learning Rate Schedule via Checkpoint Merging for LLM Pre-trainingChangxin Tian, jiapeng wang, Qian Zhao, Kunlong Chen 等ICLR 2026 · 被引用 20 次
- Model Fusion through Bayesian Optimization in Language Model Fine-TuningChaeyun Jang, Hyungi Lee, Jungtaek Kim, Juho LeeNeurIPS 2024 · 被引用 8 次
- Lifelong Test-Time Adaptation via Online Learning in Tracked Low-Dimensional SubspaceDexin Duan, Rui Xu, Peilin Liu, Fei WenNeurIPS 2025 · 被引用 7 次
相关 Paper
- Improving Neural Network Training in Low Dimensional Random BasesFrithjof Gressmann, Zach Eaton-Rosen, Carlo LuschiNeurIPS 2020 · 被引用 35 次
- Stochastic Weight Averaging in Parallel: Large-Batch Training That Generalizes WellVipul Gupta, Santiago Akle Serrano, Dennis DeCosteICLR 2020 · 被引用 78 次
- Don't Use Large Mini-batches, Use Local SGDTao Lin, Sebastian U. Stich, Kumar Kshitij Patel, Martin JaggiICLR 2020 · 被引用 462 次
- When, Where and Why to Average Weights?Niccolò Ajroldi, Antonio Orvieto, Jonas GeipingICML 2025
- Neural networks with late-phase weightsJohannes von Oswald, Seijin Kobayashi, João Sacramento, Alexander Meulemans 等ICLR 2021 · 被引用 38 次
