Learning to Boost Training by Periodic Nowcasting Near Future Weights
Jinhyeok Jang, Woo-han Yun, Won Hwa Kim, Youngwoo Yoon, Jaehong Kim, Jaeyeon Lee, ByungOk Han
Abstract
Recent complicated problems require large-scale datasets and complex model architectures, however, it is difficult to train such large networks due to high computational issues. Significant efforts have been made to make the training more efficient such as momentum, learning rate scheduling, weight regularization, and meta-learning. Based on our observations on 1) high correlation between past weights and future weights, 2) conditions for beneficial weight prediction, and 3) feasibility of weight prediction, we propose a more general framework by intermittently skipping a handful of epochs by periodically forecasting near future weights, i.e., a Weight Nowcaster Network (WNN). As an add-on module, WNN predicts the future weights to make the learning process faster regardless of tasks and architectures. Experimental results show that WNN can significantly save actual time cost for training with an additional marginal time to train WNN. We validate the generalization capability of WNN under various tasks, and demonstrate that it works well even for unseen tasks. The code and pre-trained model are available at https://github.com/jjh6297/WNN .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 599a1618-dff6-4bc3-aeeb-7ef4b28209ecCited by top-tier papers4
- Learning to Rewind via Iterative Prediction of Past Weights for Practical UnlearningJinhyeok Jang, Jaehong Kim, Chan-Hyun YounAAAI 2025 · 2 citations
- Learning from Oblivion: Predicting Knowledge-Overflowed Weights via Retrodiction of ForgettingJinhyeok Jang, Jaehong Kim, Jung Uk KimCVPR 2026
- Accelerating Training with Neuron Interaction and Nowcasting NetworksBoris Knyazev, Abhinav Moudgil, Guillaume Lajoie, Eugene Belilovsky et al.ICLR 2025
- Predictive Differential Training Guided by Training DynamicsFanqi Wang, Weisheng Tang, Landon Harris, Hairong Qi et al.ICLR 2026
Builds on5
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Parameter Prediction for Unseen Deep ArchitecturesBoris Knyazev, Michal Drozdzal, Graham W. Taylor, Adriana Romero-SorianoNeurIPS 2021 · 111 citations
- GradInit: Learning to Initialize Neural Networks for Stable and Efficient TrainingChen Zhu, Renkun Ni, Zheng Xu, Kezhi Kong et al.NeurIPS 2021 · 73 citations
Related papers
- Learning with RetrospectionXiang Deng, Zhongfei ZhangAAAI 2021 · 20 citations
- Rapid Neural Architecture Search by Learning to Generate Graphs from DatasetsHayeon Lee, Eunyoung Hyung, Sung Ju HwangICLR 2021 · 57 citations
- Prediction Confidence based Low Complexity Gradient Computation for Accelerating DNN TrainingDongyeob Shin, Geonho Kim, Joongho Jo, Jongsun ParkDAC 2020 · 14 citations
- On the Overlooked Pitfalls of Weight Decay and How to Mitigate Them: A Gradient-Norm PerspectiveZeke Xie, Zhiqiang Xu, Jingzhao Zhang, Issei Sato et al.NeurIPS 2023 · 38 citations
- Scaling Laws and Compute-Optimal Training Beyond Fixed Training DurationsAlexander Hägele, Elie Bakouch, Atli Kosson, Loubna Ben Allal et al.NeurIPS 2024 · 168 citations
