Better Training using Weight-Constrained Stochastic Dynamics
Benedict J. Leimkuhler, Tiffany J. Vlaar, Timothée Pouchon, Amos J. Storkey
摘要
We employ constraints to control the parameter space of deep neural networks throughout training. The use of customized, appropriately designed constraints can reduce the vanishing/exploding gradients problem, improve smoothness of classification boundaries, control weight magnitudes and stabilize deep neural networks, and thus enhance the robustness of training algorithms and the generalization capabilities of neural networks. We provide a general approach to efficiently incorporate constraints into a stochastic gradient Langevin framework, allowing enhanced exploration of the loss landscape. We also present specific examples of constrained training methods motivated by orthogonality preservation for weight matrices and explicit weight normalizations. Discretization schemes are provided both for the overdamped formulation of Langevin dynamics and the underdamped form, in which momenta further improve sampling efficiency. These optimization schemes can be used directly, without needing to adapt neural network architecture design choices or to modify the objective with regularization terms, and see performance improvements in classification tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Task-Driven Wavelets Using Constrained Empirical Risk MinimizationEric Marcus, Ray Sheombarsing, Jan-Jakob Sonke, Jonas TeuwenCVPR 2024
- Momentum Stiefel Optimizer, with Applications to Suitably-Orthogonal Attention, and Optimal TransportLingkai Kong, Yuqing Wang, Molei TaoICLR 2023
它引用的顶会 Paper1
相关 Paper
- Improved Analysis of Clipping Algorithms for Non-convex OptimizationBohang Zhang, Jikai Jin, Cong Fang, Liwei WangNeurIPS 2020 · 被引用 139 次
- Surprising Instabilities in Training Deep Networks and a Theoretical AnalysisYuxin Sun, Dong Lao, Ganesh Sundaramoorthi, Anthony J. YezziNeurIPS 2022 · 被引用 15 次
- Chaotic Dynamics are Intrinsic to Neural Network Training with SGDLuis Herrmann, Maximilian Granz, Tim LandgrafNeurIPS 2022 · 被引用 15 次
- On the Overlooked Pitfalls of Weight Decay and How to Mitigate Them: A Gradient-Norm PerspectiveZeke Xie, Zhiqiang Xu, Jingzhao Zhang, Issei Sato 等NeurIPS 2023 · 被引用 38 次
- Expected Gradients of Maxout Networks and Consequences to Parameter InitializationHanna Tseran, Guido MontúfarICML 2023 · 被引用 1 次
