Grounding and Enhancing Informativeness and Utility in Dataset Distillation
Shaobo Wang, Yantai Yang, Guo Chen, Peiru Li, Kaixin Li, Yufa Zhou, Zhaorun Chen, Linfeng Zhang
Abstract
Dataset Distillation (DD) seeks to create a compact dataset from a large, realworld dataset. While recent methods often rely on heuristic approaches to balance efficiency and quality, the fundamental relationship between original and synthetic data remains underexplored. This paper revisits knowledge distillationbased dataset distillation within a solid theoretical framework. We introduce the concepts of Informativeness and Utility, capturing crucial information within a sample and essential samples in the training set, respectively. Building on these principles, we define optimal dataset distillation mathematically. We then present InfoUtil, a framework that balances informativeness and utility in synthesizing the distilled dataset. InfoUtil incorporates two key components: (1) game-theoretic informativeness maximization using Shapley Value attribution to extract key information from samples, and (2) principled utility maximization by selecting globally influential samples based on Gradient Norm. These components ensure that the distilled dataset is both informative and utility-optimized. Experiments demonstrate that our method achieves a 6.1% performance improvement over the previous state-of-the-art approach on ImageNet-1K dataset using ResNet-18. (b) InfoUtil (Ours) Random Cropping & Loss Scoring (not interpretable or theoretically principled) Attribution Cropping & GradNorm Scoring (both interpretable and theoretically principled) (a) RDED Cock Brambling Otter
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0966a88d-d362-4154-af23-a44fdfb9869bBuilds on21
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 684 citations
- Dataset Condensation with Differentiable Siamese AugmentationBo Zhao, Hakan BilenICML 2021 · 390 citations
Related papers
- Learnability-Guided Diffusion for Dataset DistillationJeffrey A. Chan-Santiago, Mubarak ShahCVPR 2026 · 2 citations
- MIM4DD: Mutual Information Maximization for Dataset DistillationYuzhang Shang, Zhihang Yuan, Yan YanNeurIPS 2023 · 26 citations
- Influence-Guided Diffusion for Dataset DistillationMingyang Chen, Jiawei Du, Bo Huang, Yi Wang et al.ICLR 2025
- Diffusion Models as Dataset Distillation PriorsDuo Su, Huyu Wu, Huanran Chen, Yiming Shi et al.ICLR 2026 · 3 citations
- Breaking Class Barriers: Efficient Dataset Distillation via Inter-Class Feature CompensatorXin Zhang, Jiawei Du, Ping Liu, Joey Tianyi ZhouICLR 2025
