Generalization bounds via distillation
Daniel Hsu, Ziwei Ji, Matus Telgarsky, Lan Wang
摘要
This paper theoretically investigates the following empirical phenomenon: given a highcomplexity network with poor generalization bounds, one can distill it into a network with nearly identical predictions but low complexity and vastly smaller generalization bounds. The main contribution is an analysis showing that the original network inherits this good generalization bound from its distillation, assuming the use of well-behaved data augmentation. This bound is presented both in an abstract and in a concrete form, the latter complemented by a reduction technique to handle modern computation graphs featuring convolutional layers, fully-connected layers, and skip connections, to name a few. To round out the story, a (looser) classical uniform convergence analysis of compression is also presented, as well as a variety of experiments on cifar10 and mnist demonstrating similar generalization performance between the original network and its distillation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- A statistical perspective on distillationAditya Krishna Menon, Ankit Singh Rawat, Sashank J. Reddi, Seungyeon Kim 等ICML 2021 · 被引用 97 次
- Intrinsic Dimension, Persistent Homology and Generalization in Neural NetworksTolga Birdal, Aaron Lou, Leonidas J. Guibas, Umut SimsekliNeurIPS 2021 · 被引用 94 次
- Heavy Tails in SGD and Compressibility of Overparametrized Neural NetworksMelih Barsbey, Milad Sefidgaran, Murat A. Erdogdu, Gaël Richard 等NeurIPS 2021 · 被引用 57 次
- Quality-Agnostic Deepfake Detection with Intra-model Collaborative LearningBinh Minh Le, Simon S. WooICCV 2023 · 被引用 50 次
- Asymmetric Temperature Scaling Makes Larger Networks Teach Well AgainXin-Chun Li, Wen-Shu Fan, Shaoming Song, Yinchuan Li 等NeurIPS 2022 · 被引用 46 次
它引用的顶会 Paper6
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan 等ICLR 2020 · 被引用 705 次
- Pruning Neural Networks at Initialization: Why Are We Missing the Mark?Jonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICLR 2021 · 被引用 261 次
- In search of robust measures of generalizationGintare Karolina Dziugaite, Alexandre Drouin, Brady Neal, Nitarshan Rajkumar 等NeurIPS 2020 · 被引用 112 次
- Generalization bounds for deep convolutional neural networksPhilip M. Long, Hanie SedghiICLR 2020 · 被引用 102 次
- Sanity-Checking Pruning Methods: Random Tickets can Win the JackpotJingtong Su, Yihang Chen, Tianle Cai, Tianhao Wu 等NeurIPS 2020 · 被引用 100 次
相关 Paper
- Dataset Distillation Efficiently Encodes Low-Dimensional Representations from Gradient-Based Learning of Non-Linear TasksYuri Kinoshita, Naoki Nishikawa, Taro ToyoizumiICML 2026
- Compression based bound for non-compressed network: unified generalization error analysis of large compressible deep neural networkTaiji Suzuki, Hiroshi Abe, Tomoaki NishimuraICLR 2020 · 被引用 57 次
- Utility Boundary of Dataset Distillation: Scaling and Coverage LawsZhengquan Luo, Zhiqiang XuICML 2026
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen 等ICCV 2019 · 被引用 1,069 次
- Lipschitz Continuity Guided Knowledge DistillationYuzhang Shang, Bin Duan, Ziliang Zong, Liqiang Nie 等ICCV 2021 · 被引用 31 次
