Network-to-Network Regularization: Enforcing Occam's Razor to Improve Generalization
Rohan Ghosh, Mehul Motani
摘要
What makes a classifier have the ability to generalize? There have been a lot of important attempts to address this question, but a clear answer is still elusive. Proponents of complexity theory find that the complexity of the classifier's function space is key to deciding generalization, whereas other recent work reveals that classifiers which extract invariant feature representations are likely to generalize better. Recent theoretical and empirical studies, however, have shown that even within a classifier's function space, there can be significant differences in the ability to generalize. Specifically, empirical studies have shown that among functions which have a good training data fit, functions with lower Kolmogorov complexity (KC) are likely to generalize better, while the opposite is true for functions of higher KC. Motivated by these findings, we propose, in this work, a novel measure of complexity called Kolmogorov Growth (KG), which we use to derive new generalization error bounds that only depend on the final choice of the classification function. Guided by the bounds, we propose a novel way of regularizing neural networks by constraining the network trajectory to remain in the low KG zone during training. Minimizing KG while learning is akin to applying the Occam's razor to neural networks. The proposed approach, called network-to-network regularization, leads to clear improvements in the generalization ability of classifiers. We verify this for three popular image datasets (MNIST, CIFAR-10, CIFAR-100) across varying training data sizes. Empirical studies find that conventional training of neural networks, unlike network-to-network regularization, leads to networks of high KG and lower test accuracies. Furthermore, we present the benefits of N2N regularization in the scenario where the training data labels are noisy. Using N2N regularization, we achieve competitive performance on MNIST, CIFAR-10 and CIFAR-100 datasets with corrupted training labels, significantly improving network performance compared to standard cross-entropy baselines in most cases. These findings illustrate the many benefits obtained from imposing a function complexity prior like Kolmogorov Growth during the training process.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Symmetric Cross Entropy for Robust Learning With Noisy LabelsYisen Wang, Xingjun Ma, Zaiyi Chen, Yuan Luo 等ICCV 2019 · 被引用 1,125 次
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan 等ICLR 2020 · 被引用 705 次
- In search of robust measures of generalizationGintare Karolina Dziugaite, Alexandre Drouin, Brady Neal, Nitarshan Rajkumar 等NeurIPS 2020 · 被引用 112 次
- Generalization bounds via distillationDaniel Hsu, Ziwei Ji, Matus Telgarsky, Lan WangICLR 2021 · 被引用 14 次
- Revisiting Knowledge Distillation via Label Smoothing RegularizationLi Yuan, Francis E. H. Tay, Guilin Li, Tao Wang 等CVPR 2020
相关 Paper
- Neural Complexity MeasuresYoonho Lee, Juho Lee, Sung Ju Hwang, Eunho Yang 等NeurIPS 2020 · 被引用 13 次
- Understanding Generalization in Recurrent Neural NetworksZhuozhuo Tu, Fengxiang He, Dacheng TaoICLR 2020 · 被引用 33 次
- Measuring Model Complexity of Neural Networks with Curve Activation FunctionsXia Hu, Weiqing Liu, Jiang Bian, Jian PeiKDD 2020 · 被引用 25 次
- Feature Variance Regularization: A Simple Way to Improve the Generalizability of Neural NetworksRanran Huang, Hanbo Sun, Ji Liu, Lu Tian 等AAAI 2020 · 被引用 5 次
- On the Interpretability of Regularisation for Neural Networks Through Model Gradient SimilarityVincent Szolnoky, Viktor Andersson, Balázs Kulcsár, Rebecka JörnstenNeurIPS 2022 · 被引用 6 次
