Better Training using Weight-Constrained Stochastic Dynamics
Benedict J. Leimkuhler, Tiffany J. Vlaar, Timothée Pouchon, Amos J. Storkey
Abstract
We employ constraints to control the parameter space of deep neural networks throughout training. The use of customized, appropriately designed constraints can reduce the vanishing/exploding gradients problem, improve smoothness of classification boundaries, control weight magnitudes and stabilize deep neural networks, and thus enhance the robustness of training algorithms and the generalization capabilities of neural networks. We provide a general approach to efficiently incorporate constraints into a stochastic gradient Langevin framework, allowing enhanced exploration of the loss landscape. We also present specific examples of constrained training methods motivated by orthogonality preservation for weight matrices and explicit weight normalizations. Discretization schemes are provided both for the overdamped formulation of Langevin dynamics and the underdamped form, in which momenta further improve sampling efficiency. These optimization schemes can be used directly, without needing to adapt neural network architecture design choices or to modify the objective with regularization terms, and see performance improvements in classification tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Task-Driven Wavelets Using Constrained Empirical Risk MinimizationEric Marcus, Ray Sheombarsing, Jan-Jakob Sonke, Jonas TeuwenCVPR 2024
- Momentum Stiefel Optimizer, with Applications to Suitably-Orthogonal Attention, and Optimal TransportLingkai Kong, Yuqing Wang, Molei TaoICLR 2023
Builds on1
Related papers
- Improved Analysis of Clipping Algorithms for Non-convex OptimizationBohang Zhang, Jikai Jin, Cong Fang, Liwei WangNeurIPS 2020 · 139 citations
- Surprising Instabilities in Training Deep Networks and a Theoretical AnalysisYuxin Sun, Dong Lao, Ganesh Sundaramoorthi, Anthony J. YezziNeurIPS 2022 · 15 citations
- Chaotic Dynamics are Intrinsic to Neural Network Training with SGDLuis Herrmann, Maximilian Granz, Tim LandgrafNeurIPS 2022 · 15 citations
- On the Overlooked Pitfalls of Weight Decay and How to Mitigate Them: A Gradient-Norm PerspectiveZeke Xie, Zhiqiang Xu, Jingzhao Zhang, Issei Sato et al.NeurIPS 2023 · 38 citations
- Expected Gradients of Maxout Networks and Consequences to Parameter InitializationHanna Tseran, Guido MontúfarICML 2023 · 1 citation
