Predictive Differential Training Guided by Training Dynamics
Fanqi Wang, Weisheng Tang, Landon Harris, Hairong Qi, Dan Wilson, Igor Mezic
Abstract
This paper centers around a novel concept proposed recently by researchers from the control community where the training process of a deep neural network can be considered a nonlinear dynamical system acting upon the high-dimensional weight space. Koopman operator theory (KOT), a data-driven dynamical system analysis framework, can then be deployed to discover the otherwise non-intuitive training dynamics. Taking advantage of the predictive power of KOT, the time-consuming Stochastic Gradient Descent (SGD) iterations can be then bypassed by directly predicting network weights a few epochs later. This "predictive training" framework, however, often suffers from gradient explosion especially for more extensive and complex models. In this paper, we incorporate the idea of "differential learning" into the predictive training framework and propose the so-called "predictive differential training" (PDT) for accelerated learning even for complex network structures. The key contribution is the design of an effective masking strategy based on a dynamic consistency analysis, which selects only those predicted weights whose local training dynamics align with the global dynamics. We refer to these predicted weights as high-fidelity predictions. PDT also includes the design of an acceleration scheduler to adjust the prediction interval and rectify deviations from off-predictions. We demonstrate that PDT can be seamlessly integrated as a plug-in with a diverse array of existing optimizers (SGD, Adam, RMSprop, LAMB, etc.). The experimental results show consistent performance improvement across different network architectures and various datasets, in terms of faster convergence and reduced training time (10-40%) to achieve the baseline's best loss, while maintaining (if not improving) final model accuracy. As the idiom goes, a rising tide lifts all boats; in our context, a subset of high-fidelity predicted weights can accelerate the training of the entire network! * Equal contribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on6
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu et al.ICLR 2020 · 1,170 citations
- Optimizing Neural Networks via Koopman Operator TheoryAkshunna S. Dogra, William T. RedmanNeurIPS 2020 · 65 citations
- An Operator Theoretic View On Pruning Deep Neural NetworksWilliam T. Redman, Maria Fonoberova, Ryan Mohr, Yannis G. Kevrekidis et al.ICLR 2022 · 21 citations
- Identifying Equivalent Training DynamicsWilliam T. Redman, Juan M. Bello-Rivas, Maria Fonoberova, Ryan Mohr et al.NeurIPS 2024 · 15 citations
- Learning to Boost Training by Periodic Nowcasting Near Future WeightsJinhyeok Jang, Woo-han Yun, Won Hwa Kim, Youngwoo Yoon et al.ICML 2023 · 6 citations
Related papers
- Second-Order Neural ODE OptimizerGuan-Horng Liu, Tianrong Chen, Evangelos A. TheodorouNeurIPS 2021 · 20 citations
- Dynamic Game Theoretic Neural OptimizerGuan-Horng Liu, Tianrong Chen, Evangelos A. TheodorouICML 2021 · 6 citations
- DDPNOpt: Differential Dynamic Programming Neural OptimizerGuan-Horng Liu, Tianrong Chen, Evangelos A. TheodorouICLR 2021 · 7 citations
- QPKO: Differentiable QP-Embedded Deep Koopman Framework for Modeling Nonlinear SystemsRunze Tian, Peng KouICML 2026
- Advancing Dynamic Sparse Training by Exploring Optimization OpportunitiesJie Ji, Gen Li, Lu Yin, Minghai Qin et al.ICML 2024 · 10 citations
