Slimmable Dataset Condensation
Songhua Liu, Jingwen Ye, Runpeng Yu, Xinchao Wang
Abstract
Dataset distillation, also known as dataset condensation, aims to compress a large dataset into a compact synthetic one. Existing methods perform dataset condensation by assuming a fixed storage or transmission budget. When the budget changes, however, they have to repeat the synthesizing process with access to original datasets, which is highly cumbersome if not infeasible at all. In this paper, we explore the problem of slimmable dataset condensation, to extract a smaller synthetic dataset given only previous condensation results. We first study the limitations of existing dataset condensation algorithms on such a successive compression setting and identify two key factors: (1) the inconsistency of neural networks over different compression times and (2) the underdetermined solution space for synthetic data. Accordingly, we propose a novel training objective for slimmable dataset condensation to explicitly account for both factors. Moreover, synthetic datasets in our method adopt a significance-aware parameterization. Theoretical derivation indicates that an upper-bounded error can be achieved by discarding the minor components without training. Alternatively, if training is allowed, this strategy can serve as a strong initialization that enables a fast convergence. Extensive comparisons and ablations demonstrate the superiority of the proposed solution over existing methods on multiple benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 73a9acb6-12f0-4ba3-a538-584a8de4577dCited by top-tier papers23
- Towards Lossless Dataset Distillation via Difficulty-Aligned Trajectory MatchingZiyao Guo, Kai Wang, George Cazenavette, Hui Li et al.ICLR 2024 · 142 citations
- Diffusion Model as Representation LearnerXingyi Yang, Xinchao WangICCV 2023 · 100 citations
- SG-Former: Self-guided Transformer with Evolving Token ReallocationSucheng Ren, Xingyi Yang, Songhua Liu, Xinchao WangICCV 2023 · 70 citations
- Evaluation and Improvement of Interpretability for Self-Explainable Part-Prototype NetworksQihan Huang, Mengqi Xue, Wenqi Huang, Haofei Zhang et al.ICCV 2023 · 47 citations
- Navigating Complexity: Toward Lossless Graph Condensation via Expanding Window MatchingYuchen Zhang, Tianle Zhang, Kai Wang, Ziyao Guo et al.ICML 2024 · 38 citations
Builds on23
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 684 citations
- Dataset Condensation with Differentiable Siamese AugmentationBo Zhao, Hakan BilenICML 2021 · 390 citations
- Shunted Self-Attention via Multi-Scale Token AggregationSucheng Ren, Daquan Zhou, Shengfeng He, Jiashi Feng et al.CVPR 2022 · 326 citations
- Dataset Distillation with Infinitely Wide Convolutional NetworksTimothy Nguyen, Roman Novak, Lechao Xiao, Jaehoon LeeNeurIPS 2021 · 313 citations
Related papers
- Dataset Condensation via Efficient Synthetic-Data ParameterizationJang-Hyun Kim, Jinuk Kim, Seong Joon Oh, Sangdoo Yun et al.ICML 2022 · 234 citations
- CondTSF: One-line Plugin of Dataset Condensation for Time Series ForecastingJianrong Ding, Zhanyu Liu, Guanjie Zheng, Haiming Jin et al.NeurIPS 2024 · 8 citations
- Frequency Domain-Based Dataset DistillationDongHyeok Shin, Seungjae Shin, Il-Chul MoonNeurIPS 2023 · 39 citations
- DREAM: Efficient Dataset Distillation by Representative MatchingYanqing Liu, Jianyang Gu, Kai Wang, Zheng Zhu et al.ICCV 2023 · 114 citations
- Improved Distribution Matching for Dataset CondensationGanlong Zhao, Guanbin Li, Yipeng Qin, Yizhou YuCVPR 2023
