Network-to-Network Regularization: Enforcing Occam's Razor to Improve Generalization
Rohan Ghosh, Mehul Motani
Abstract
What makes a classifier have the ability to generalize? There have been a lot of important attempts to address this question, but a clear answer is still elusive. Proponents of complexity theory find that the complexity of the classifier's function space is key to deciding generalization, whereas other recent work reveals that classifiers which extract invariant feature representations are likely to generalize better. Recent theoretical and empirical studies, however, have shown that even within a classifier's function space, there can be significant differences in the ability to generalize. Specifically, empirical studies have shown that among functions which have a good training data fit, functions with lower Kolmogorov complexity (KC) are likely to generalize better, while the opposite is true for functions of higher KC. Motivated by these findings, we propose, in this work, a novel measure of complexity called Kolmogorov Growth (KG), which we use to derive new generalization error bounds that only depend on the final choice of the classification function. Guided by the bounds, we propose a novel way of regularizing neural networks by constraining the network trajectory to remain in the low KG zone during training. Minimizing KG while learning is akin to applying the Occam's razor to neural networks. The proposed approach, called network-to-network regularization, leads to clear improvements in the generalization ability of classifiers. We verify this for three popular image datasets (MNIST, CIFAR-10, CIFAR-100) across varying training data sizes. Empirical studies find that conventional training of neural networks, unlike network-to-network regularization, leads to networks of high KG and lower test accuracies. Furthermore, we present the benefits of N2N regularization in the scenario where the training data labels are noisy. Using N2N regularization, we achieve competitive performance on MNIST, CIFAR-10 and CIFAR-100 datasets with corrupted training labels, significantly improving network performance compared to standard cross-entropy baselines in most cases. These findings illustrate the many benefits obtained from imposing a function complexity prior like Kolmogorov Growth during the training process.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4da6c457-775c-43c6-9f54-5a6937ced78dBuilds on5
- Symmetric Cross Entropy for Robust Learning With Noisy LabelsYisen Wang, Xingjun Ma, Zaiyi Chen, Yuan Luo et al.ICCV 2019 · 1,125 citations
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan et al.ICLR 2020 · 705 citations
- In search of robust measures of generalizationGintare Karolina Dziugaite, Alexandre Drouin, Brady Neal, Nitarshan Rajkumar et al.NeurIPS 2020 · 112 citations
- Generalization bounds via distillationDaniel Hsu, Ziwei Ji, Matus Telgarsky, Lan WangICLR 2021 · 14 citations
- Revisiting Knowledge Distillation via Label Smoothing RegularizationLi Yuan, Francis E. H. Tay, Guilin Li, Tao Wang et al.CVPR 2020
Related papers
- Neural Complexity MeasuresYoonho Lee, Juho Lee, Sung Ju Hwang, Eunho Yang et al.NeurIPS 2020 · 13 citations
- Understanding Generalization in Recurrent Neural NetworksZhuozhuo Tu, Fengxiang He, Dacheng TaoICLR 2020 · 33 citations
- Measuring Model Complexity of Neural Networks with Curve Activation FunctionsXia Hu, Weiqing Liu, Jiang Bian, Jian PeiKDD 2020 · 25 citations
- Feature Variance Regularization: A Simple Way to Improve the Generalizability of Neural NetworksRanran Huang, Hanbo Sun, Ji Liu, Lu Tian et al.AAAI 2020 · 5 citations
- On the Interpretability of Regularisation for Neural Networks Through Model Gradient SimilarityVincent Szolnoky, Viktor Andersson, Balázs Kulcsár, Rebecka JörnstenNeurIPS 2022 · 6 citations
