Rethinking Data Distillation: Do Not Overlook Calibration
Dongyao Zhu, Yanbo Fang, Bowen Lei, Yiqun Xie, Dongkuan Xu, Jie Zhang, Ruqi Zhang
摘要
Neural networks trained on distilled data often produce over-confident output and require correction by calibration methods. Existing calibration methods such as temperature scaling and mixup work well for networks trained on original large-scale data. However, we find that these methods fail to calibrate networks trained on data distilled from large source datasets. In this paper, we show that distilled data lead to networks that are not calibratable due to (i) a more concentrated distribution of the maximum logits and (ii) the loss of information that is semantically meaningful but unrelated to classification tasks. To address this problem, we propose Masked Temperature Scaling (MTS) and Masked Distillation Training (MDT) which mitigate the limitations of distilled data and achieve better calibration results while maintaining the efficiency of dataset distillation. Our code is available upon request.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Revisit the Essence of Distilling Knowledge through CalibrationWen-Shu Fan, Su Lu, Xin-Chun Li, De-Chuan Zhan 等ICML 2024 · 被引用 8 次
- Calibration Bottleneck: Over-compressed Representations are Less CalibratableDeng-Bao Wang, Min-Ling ZhangICML 2024 · 被引用 7 次
- Optimizing Distributional Geometry Alignment with Optimal Transport for Generative Dataset DistillationXiao Cui, Yulei Qin, Wengang Zhou, Hongsheng Li 等NeurIPS 2025 · 被引用 5 次
- Towards Understanding The Calibration Benefits of Sharpness-Aware MinimizationChengli Tan, Yubo Zhou, Haishan Ye, Guang Dai 等ICLR 2026 · 被引用 3 次
- Rethinking Long-tailed Dataset Distillation: A Uni-Level Framework with Unbiased Recovery and RelabelingXiao Cui, Yulei Qin, Xinyue Li, Wengang Zhou 等AAAI 2026 · 被引用 1 次
它引用的顶会 Paper20
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 被引用 684 次
- Calibrating Deep Neural Networks using Focal LossJishnu Mukhoti, Viveka Kulharia, Amartya Sanyal, Stuart Golodetz 等NeurIPS 2020 · 被引用 674 次
- Dataset Condensation with Differentiable Siamese AugmentationBo Zhao, Hakan BilenICML 2021 · 被引用 390 次
- Mitigating Neural Network Overconfidence with Logit NormalizationHongxin Wei, Renchunzi Xie, Hao Cheng, Lei Feng 等ICML 2022 · 被引用 386 次
- Dataset Distillation with Infinitely Wide Convolutional NetworksTimothy Nguyen, Roman Novak, Lechao Xiao, Jaehoon LeeNeurIPS 2021 · 被引用 313 次
相关 Paper
- On the Limitations of Temperature Scaling for Distributions with OverlapsMuthu Chidambaram, Rong GeICLR 2024 · 被引用 11 次
- Ensemble Distillation for Structured Prediction: Calibrated, Accurate, Fast - Choose ThreeSteven Reich, David Mueller, Nicholas AndrewsEMNLP 2020 · 被引用 1 次
- Uncertainty Quantification and Deep EnsemblesRahul Rahaman, Alexandre H. ThiéryNeurIPS 2021 · 被引用 250 次
- Transfer Knowledge from Head to Tail: Uncertainty Calibration under Long-tailed DistributionJiahao Chen, Bing SuCVPR 2023
- Asymmetric Temperature Scaling Makes Larger Networks Teach Well AgainXin-Chun Li, Wen-Shu Fan, Shaoming Song, Yinchuan Li 等NeurIPS 2022 · 被引用 46 次
