Rethinking Data Distillation: Do Not Overlook Calibration
Dongyao Zhu, Yanbo Fang, Bowen Lei, Yiqun Xie, Dongkuan Xu, Jie Zhang, Ruqi Zhang
Abstract
Neural networks trained on distilled data often produce over-confident output and require correction by calibration methods. Existing calibration methods such as temperature scaling and mixup work well for networks trained on original large-scale data. However, we find that these methods fail to calibrate networks trained on data distilled from large source datasets. In this paper, we show that distilled data lead to networks that are not calibratable due to (i) a more concentrated distribution of the maximum logits and (ii) the loss of information that is semantically meaningful but unrelated to classification tasks. To address this problem, we propose Masked Temperature Scaling (MTS) and Masked Distillation Training (MDT) which mitigate the limitations of distilled data and achieve better calibration results while maintaining the efficiency of dataset distillation. Our code is available upon request.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f22837a1-6313-42db-a63a-344f0705f7a4Cited by top-tier papers6
- Revisit the Essence of Distilling Knowledge through CalibrationWen-Shu Fan, Su Lu, Xin-Chun Li, De-Chuan Zhan et al.ICML 2024 · 8 citations
- Calibration Bottleneck: Over-compressed Representations are Less CalibratableDeng-Bao Wang, Min-Ling ZhangICML 2024 · 7 citations
- Optimizing Distributional Geometry Alignment with Optimal Transport for Generative Dataset DistillationXiao Cui, Yulei Qin, Wengang Zhou, Hongsheng Li et al.NeurIPS 2025 · 5 citations
- Towards Understanding The Calibration Benefits of Sharpness-Aware MinimizationChengli Tan, Yubo Zhou, Haishan Ye, Guang Dai et al.ICLR 2026 · 3 citations
- Rethinking Long-tailed Dataset Distillation: A Uni-Level Framework with Unbiased Recovery and RelabelingXiao Cui, Yulei Qin, Xinyue Li, Wengang Zhou et al.AAAI 2026 · 1 citation
Builds on20
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 684 citations
- Calibrating Deep Neural Networks using Focal LossJishnu Mukhoti, Viveka Kulharia, Amartya Sanyal, Stuart Golodetz et al.NeurIPS 2020 · 674 citations
- Dataset Condensation with Differentiable Siamese AugmentationBo Zhao, Hakan BilenICML 2021 · 390 citations
- Mitigating Neural Network Overconfidence with Logit NormalizationHongxin Wei, Renchunzi Xie, Hao Cheng, Lei Feng et al.ICML 2022 · 386 citations
- Dataset Distillation with Infinitely Wide Convolutional NetworksTimothy Nguyen, Roman Novak, Lechao Xiao, Jaehoon LeeNeurIPS 2021 · 313 citations
Related papers
- On the Limitations of Temperature Scaling for Distributions with OverlapsMuthu Chidambaram, Rong GeICLR 2024 · 11 citations
- Ensemble Distillation for Structured Prediction: Calibrated, Accurate, Fast - Choose ThreeSteven Reich, David Mueller, Nicholas AndrewsEMNLP 2020 · 1 citation
- Uncertainty Quantification and Deep EnsemblesRahul Rahaman, Alexandre H. ThiéryNeurIPS 2021 · 250 citations
- Transfer Knowledge from Head to Tail: Uncertainty Calibration under Long-tailed DistributionJiahao Chen, Bing SuCVPR 2023
- Asymmetric Temperature Scaling Makes Larger Networks Teach Well AgainXin-Chun Li, Wen-Shu Fan, Shaoming Song, Yinchuan Li et al.NeurIPS 2022 · 46 citations
