What Makes a "Good" Data Augmentation in Knowledge Distillation - A Statistical Perspective
Huan Wang, Suhas Lohit, Michael J. Jones, Yun Fu
摘要
Good-DA-in-KD Teacher (fixed) Student Raw input 𝑥 ! Input 𝑥 Standard DA Stronger DA Input 𝑥' KL Div. loss How to define "stronger" DA? 0.00425 0.00450 0.00475 0.00500 0.00525 0.00550 T. stddev 1.00 1.05 1.10 S. test loss Pearson: 0.9581 (p-value: 0.00%) Spearman: 0.9667 (p-value: 0.00%) Kendall: 0.8889 (p-value: 0.02%) wrn_40_2/wrn_16_2, CIFAR100 Identity Flip Crop+Flip Cutout AutoAugment Mixup CutMix CutmixPick (S. ent.) CutmixPick (T. ent.) (a) Apply additional stronger DA in KD (b) S. test loss vs. T. stddev with different DA schemes Figure 1: (a) Illustration of applying a stronger data augmentation (DA) in addition to the standard DA (random crop and flip) in knowledge distillation (KD). We ask: What makes a "good" DA when it is applied to KD in the manner of (a)? (b) We present a proven proposition (Proposition 3.1) to answer this question rigorously, along with a practical metric to evaluate the "goodness" of a DA. The proposed metric is called stddev of teacher's mean probability (shorted as T. stddev). As seen in (b), there is a strong positive correlation (p-value < 5% is typically considered statistically significant) between the student's test loss (S. test loss) and T. stddev, showing that T. stddev well captures the "goodness" of different DA schemes in KD. The most striking fact from this plot may be: T. stddev is purely calculated with the teacher (no any student used) while it can "predict" the relative order of the student's performance, implying the "goodness" of DA in KD probably is student-invariant.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Binarized Spectral Compressive ImagingYuanhao Cai, Yuxin Zheng, Jing Lin, Xin Yuan 等NeurIPS 2023 · 被引用 50 次
- Binarized Diffusion Model for Image Super-ResolutionZheng Chen, Haotong Qin, Yong Guo, Xiongfei Su 等NeurIPS 2024 · 被引用 36 次
- On student-teacher deviations in distillation: does it pay to disobey?Vaishnavh Nagarajan, Aditya Krishna Menon, Srinadh Bhojanapalli, Hossein Mobahi 等NeurIPS 2023 · 被引用 25 次
- TexQ: Zero-shot Network Quantization with Texture Feature Distribution CalibrationXinrui Chen, Yizhi Wang, Renao Yan, Yiqing Liu 等NeurIPS 2023 · 被引用 24 次
- Learning Better with Less: Effective Augmentation for Sample-Efficient Visual Reinforcement LearningGuozheng Ma, Linrui Zhang, Haoyu Wang, Lu Li 等NeurIPS 2023 · 被引用 24 次
它引用的顶会 Paper12
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 被引用 1,305 次
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 被引用 1,214 次
- On the Efficacy of Knowledge DistillationJang Hyun Cho, Bharath HariharanICCV 2019 · 被引用 741 次
- A Comprehensive Overhaul of Feature DistillationByeongho Heo, Jeesoo Kim, Sangdoo Yun, Hyojin Park 等ICCV 2019 · 被引用 727 次
相关 Paper
- Differentiable JPEG-based Input Perturbation for Knowledge Distillation Amplification via Conditional Mutual Information MaximizationSIYU CHEN, Kaixiang Zheng, Ahmed H. Salamah, EN-HUI YANGICLR 2026
- Knowledge Distillation Based on Transformed Teacher MatchingKaixiang Zheng, En-Hui YangICLR 2024 · 被引用 40 次
- Logit Standardization in Knowledge DistillationShangquan Sun, Wenqi Ren, Jingzhi Li, Rui Wang 等CVPR 2024 · 被引用 183 次
- Teacher as a Lenient Expert: Teacher-Agnostic Data-Free Knowledge DistillationHyunjune Shin, Dong-Wan ChoiAAAI 2024 · 被引用 8 次
- AugKD: Ingenious Augmentations Empower Knowledge Distillation for Image Super-ResolutionYun Zhang, Wei Li, Simiao Li, Hanting Chen 等ICLR 2025
