Improving Deep Learning Optimization through Constrained Parameter Regularization
Jörg K. H. Franke, Michael Hefenbrock, Gregor Köhler, Frank Hutter
摘要
Regularization is a critical component in deep learning. The most commonly used approach, weight decay, applies a constant penalty coefficient uniformly across all parameters. This may be overly restrictive for some parameters, while insufficient for others. To address this, we present Constrained Parameter Regularization (CPR) as an alternative to traditional weight decay. Unlike the uniform application of a single penalty, CPR enforces an upper bound on a statistical measure, such as the L2-norm, of individual parameter matrices. Consequently, learning becomes a constraint optimization problem, which we tackle using an adaptation of the augmented Lagrangian method. CPR introduces only a minor runtime overhead and only requires setting an upper bound. We propose simple yet efficient mechanisms for initializing this bound, making CPR rely on no hyperparameter or one, akin to weight decay. Our empirical studies on computer vision and language modeling tasks demonstrate CPR's effectiveness. The results show that CPR can outperform traditional weight decay and increase performance in pre-training and fine-tuning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Learning in Compact Spaces with Approximately Normalized TransformerJörg K. H. Franke, Urs Spiegelhalter, Marianna Nezhurina, Jenia Jitsev 等NeurIPS 2025 · 被引用 5 次
- nGPT: Normalized Transformer with Representation Learning on the HypersphereIlya Loshchilov, Cheng-Ping Hsieh, Simeng Sun, Boris GinsburgICLR 2025
- Power-Constrained Printed Neuromorphic Hardware TrainingTara Gheshlaghi, Haibin Zhao, Priyanjana Pal, Michael Hefenbrock 等DAC 2025
它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
相关 Paper
- Better Training using Weight-Constrained Stochastic DynamicsBenedict J. Leimkuhler, Tiffany J. Vlaar, Timothée Pouchon, Amos J. StorkeyICML 2021 · 被引用 11 次
- On the Overlooked Pitfalls of Weight Decay and How to Mitigate Them: A Gradient-Norm PerspectiveZeke Xie, Zhiqiang Xu, Jingzhao Zhang, Issei Sato 等NeurIPS 2023 · 被引用 38 次
- Adaptive Budget Allocation for Parameter-Efficient Fine-TuningQingru Zhang, Minshuo Chen, Alexander Bukharin, Pengcheng He 等ICLR 2023 · 被引用 32 次
- A Unified DNN Weight Pruning Framework Using Reweighted Optimization MethodsTianyun Zhang, Xiaolong Ma, Zheng Zhan, Shanglin Zhou 等DAC 2021 · 被引用 25 次
- AlphaDecay: Module-wise Weight Decay for Heavy-Tailed Balancing in LLMsDi He, Songjun Tu, Ajay Jaiswal, Li Shen 等NeurIPS 2025 · 被引用 14 次
