Regularizing activations in neural networks via distribution matching with the Wasserstein metric
Taejong Joo, Donggu Kang, Byunghoon Kim
摘要
Regularization and normalization have become indispensable components in training deep neural networks, resulting in faster training and improved generalization performance. We propose the projected error function regularization loss (PER) that encourages activations to follow the standard normal distribution. PER randomly projects activations onto one-dimensional space and computes the regularization loss in the projected space. PER is similar to the Pseudo-Huber loss in the projected space, thus taking advantage of both and regularization losses. Besides, PER can capture the interaction between hidden units by projection vector drawn from a unit sphere. By doing so, PER minimizes the upper bound of the Wasserstein distance of order one between an empirical distribution of activations and the standard normal distribution. To the best of the authors' knowledge, this is the first work to regularize activations via distribution matching in the probability distribution space. We evaluate the proposed method on the image classification task and the word-level language modeling task.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- On the Importance of Gaussianizing RepresentationsDaniel Eftekhari, Vardan PapyanICML 2025
- Learning with Noisy Labels via Sparse RegularizationXiong Zhou, Xianming Liu, Chenyang Wang, Deming Zhai 等ICCV 2021 · 被引用 77 次
- Exploiting Space Folding by Neural NetworksMichal Lewandowski, Raphael Pisoni, Bernhard Heinzl, Bernhard Alois MoserAAAI 2026
- Moment- and Power-Spectrum-Based Gaussianity Regularization for Text-to-Image ModelsJisung Hwang, Jaihoon Kim, Minhyuk SungNeurIPS 2025 · 被引用 2 次
- Learning Regularizer for Monocular Depth Estimation with Adversarial GuidanceGuibao Shen, Yingkui Zhang, Jialu Li, Mingqiang Wei 等ACM MM 2021 · 被引用 7 次
