Grounding and Enhancing Informativeness and Utility in Dataset Distillation
Shaobo Wang, Yantai Yang, Guo Chen, Peiru Li, Kaixin Li, Yufa Zhou, Zhaorun Chen, Linfeng Zhang
摘要
Dataset Distillation (DD) seeks to create a compact dataset from a large, realworld dataset. While recent methods often rely on heuristic approaches to balance efficiency and quality, the fundamental relationship between original and synthetic data remains underexplored. This paper revisits knowledge distillationbased dataset distillation within a solid theoretical framework. We introduce the concepts of Informativeness and Utility, capturing crucial information within a sample and essential samples in the training set, respectively. Building on these principles, we define optimal dataset distillation mathematically. We then present InfoUtil, a framework that balances informativeness and utility in synthesizing the distilled dataset. InfoUtil incorporates two key components: (1) game-theoretic informativeness maximization using Shapley Value attribution to extract key information from samples, and (2) principled utility maximization by selecting globally influential samples based on Gradient Norm. These components ensure that the distilled dataset is both informative and utility-optimized. Experiments demonstrate that our method achieves a 6.1% performance improvement over the previous state-of-the-art approach on ImageNet-1K dataset using ResNet-18. (b) InfoUtil (Ours) Random Cropping & Loss Scoring (not interpretable or theoretically principled) Attribution Cropping & GradNorm Scoring (both interpretable and theoretically principled) (a) RDED Cock Brambling Otter
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le 等ICCV 2019 · 被引用 9,163 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 被引用 684 次
- Dataset Condensation with Differentiable Siamese AugmentationBo Zhao, Hakan BilenICML 2021 · 被引用 390 次
相关 Paper
- Learnability-Guided Diffusion for Dataset DistillationJeffrey A. Chan-Santiago, Mubarak ShahCVPR 2026 · 被引用 2 次
- MIM4DD: Mutual Information Maximization for Dataset DistillationYuzhang Shang, Zhihang Yuan, Yan YanNeurIPS 2023 · 被引用 26 次
- Influence-Guided Diffusion for Dataset DistillationMingyang Chen, Jiawei Du, Bo Huang, Yi Wang 等ICLR 2025
- Diffusion Models as Dataset Distillation PriorsDuo Su, Huyu Wu, Huanran Chen, Yiming Shi 等ICLR 2026 · 被引用 3 次
- Breaking Class Barriers: Efficient Dataset Distillation via Inter-Class Feature CompensatorXin Zhang, Jiawei Du, Ping Liu, Joey Tianyi ZhouICLR 2025
