Transferring Learning Trajectories of Neural Networks
Daiki Chijiwa
摘要
Training deep neural networks (DNNs) is computationally expensive, which is problematic especially when performing duplicated or similar training runs in model ensemble or fine-tuning pre-trained models, for example. Once we have trained one DNN on some dataset, we have its learning trajectory (i.e., a sequence of intermediate parameters during training) which may potentially contain useful information for learning the dataset. However, there has been no attempt to utilize such information of a given learning trajectory for another training. In this paper, we formulate the problem of "transferring" a given learning trajectory from one initial parameter to another one (named learning transfer problem) and derive the first algorithm to approximately solve it by matching gradients successively along the trajectory via permutation symmetry. We empirically show that the transferred parameters achieve non-trivial accuracy before any direct training, and can be trained significantly faster than training from scratch.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Portable Reward Tuning: Towards Reusable Fine-Tuning across Different Pretrained ModelsDaiki Chijiwa, Taku Hasegawa, Kyosuke Nishida, Kuniko Saito 等ICML 2025
- Transferable Model-agnostic Vision-Language Model Adaptation for Efficient Weak-to-Strong GeneralizationJihwan Park, Taehoon Song, Sanghyeok Lee, Miso Choi 等AAAI 2026
它引用的顶会 Paper23
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs 等ICML 2022 · 被引用 1,464 次
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 被引用 750 次
- Merging Models with Fisher-Weighted AveragingMichael Matena, Colin RaffelNeurIPS 2022 · 被引用 741 次
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 被引用 684 次
- The Role of Permutation Invariance in Linear Mode Connectivity of Neural NetworksRahim Entezari, Hanie Sedghi, Olga Saukh, Behnam NeyshaburICLR 2022 · 被引用 301 次
相关 Paper
- ModelKeeper: Accelerating DNN Training via Automated Training WarmupFan Lai, Yinwei Dai, Harsha V. Madhyastha, Mosharaf ChowdhuryNSDI 2023 · 被引用 31 次
- Multirate Training of Neural NetworksTiffany J. Vlaar, Benedict J. LeimkuhlerICML 2022 · 被引用 6 次
- Understanding the Mechanisms of Fast Hyperparameter TransferNikhil Ghosh, Denny Wu, Alberto BiettiICLR 2026 · 被引用 8 次
- ADA-GP: Accelerating DNN Training By Adaptive Gradient PredictionVahid Janfaza, Shantanu Mandal, Farabi Mahmud, Abdullah MuzahidMICRO 2023 · 被引用 3 次
- What Do Neural Networks Learn When Trained With Random Labels?Hartmut Maennel, Ibrahim M. Alabdulmohsin, Ilya O. Tolstikhin, Robert J. N. Baldock 等NeurIPS 2020 · 被引用 99 次
