Learning from Oblivion: Predicting Knowledge-Overflowed Weights via Retrodiction of Forgetting
Jinhyeok Jang, Jaehong Kim, Jung Uk Kim
摘要
Pre-trained weights have become a cornerstone of modern deep learning, enabling efficient knowledge transfer and improving downstream task performance, especially in data-scarce scenarios. However, a fundamental question remains: how can we obtain better pre-trained weights that encapsulate more knowledge beyond the given dataset? In this work, we introduce KNowledge-Overflowed Weights (KNOW) prediction, a novel strategy that leverages structured forgetting and its inversion to synthesize knowledge-enriched weights. Our key insight is that sequential fine-tuning on progressively downsized datasets induces a structured forgetting process, which can be modeled and reversed to recover knowledge as if trained on a larger dataset. We construct a dataset of weight transitions governed by this controlled forgetting and employ meta-learning to model weight prediction effectively. Specifically, our KNowledge-Overflowed Weights Nowcaster (KNOWN) acts as a hyper-model that learns the general evolution of weights and predicts enhanced weights with improved generalization. Extensive experiments across diverse datasets and architectures demonstrate that KNOW prediction consistently outperforms Naive fine-tuning and simple weight prediction, leading to superior downstream performance. Our work provides a new perspective on reinterpreting forgetting dynamics to push the limits of knowledge transfer. The code and pre-trained model are available at https://github.com/jjh6297/KNOW
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs 等ICML 2022 · 被引用 1,464 次
- Rethinking ImageNet Pre-TrainingKaiming He, Ross B. Girshick, Piotr DollárICCV 2019 · 被引用 1,188 次
- Merging Models with Fisher-Weighted AveragingMichael Matena, Colin RaffelNeurIPS 2022 · 被引用 741 次
- Patching open-vocabulary models by interpolating weightsGabriel Ilharco, Mitchell Wortsman, Samir Yitzhak Gadre, Shuran Song 等NeurIPS 2022 · 被引用 230 次
- Overcoming Catastrophic Forgetting in Incremental Few-Shot Learning by Finding Flat MinimaGuangyuan Shi, Jiaxin Chen, Wenlong Zhang, Li-Ming Zhan 等NeurIPS 2021 · 被引用 229 次
相关 Paper
- Learning to Boost Training by Periodic Nowcasting Near Future WeightsJinhyeok Jang, Woo-han Yun, Won Hwa Kim, Youngwoo Yoon 等ICML 2023 · 被引用 6 次
- Optimizing Reusable Knowledge for Continual Learning via MetalearningJulio Hurtado, Alain Raymond-Saez, Alvaro SotoNeurIPS 2021 · 被引用 47 次
- Graceful Forgetting in Generative Language ModelsChunyang Jiang, Chi-Min Chan, Yiyang Cai, Yulong Liu 等EMNLP 2025
- On Local Overfitting and Forgetting in Deep Neural NetworksUri Stern, Tomer Yaacoby, Daphna WeinshallAAAI 2025 · 被引用 2 次
- Upweighting Easy Samples in Fine-Tuning Mitigates ForgettingSunny Sanyal, Hayden Prairie, Rudrajit Das, Ali Kavis 等ICML 2025
