Sparse Parameterization for Epitomic Dataset Distillation
Xing Wei, Anjia Cao, Funing Yang, Zhiheng Ma
Abstract
The success of deep learning relies heavily on large and diverse datasets, but the storage, preprocessing, and training of such data present significant challenges. To address these challenges, dataset distillation techniques have been proposed to obtain smaller synthetic datasets that capture the essential information of the originals. In this paper, we introduce a Sparse Parameterization for Epitomic datasEt Distilla-tion (SPEED) framework, which leverages the concept of dictionary learning and sparse coding to distill epitomes that represent pivotal information of the dataset. SPEED prioritizes proper parameterization of the synthetic dataset and introduces techniques to capture spatial redundancy within and between synthetic images. We propose Spatial-Agnostic Epitomic Tokens (SAETs) and Sparse Coding Matrices (SCMs) to efficiently represent and select significant features. Additionally, we build a Feature-Recurrent Network (FReeNet) to generate hierarchical features with high compression and storage efficiency. Experimental results demonstrate the superiority of SPEED in handling high-resolution datasets, achieving state-of-the-art performance on multiple benchmarks and downstream applications. Our framework is compatible with a variety of dataset matching approaches, generally enhancing their performance. This work highlights the importance of proper parameterization in epitomic dataset distillation and opens avenues for efficient representation learning. Source code is available at https://github.com/MIV-XJTU/SPEED .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 23f4d607-25dc-4682-993c-17eb79f1633eCited by top-tier papers17
- Evolving Parameterized Prompt Memory for Continual LearningMuhammad Rifki Kurniawan, Xiang Song, Zhiheng Ma, Yuhang He et al.AAAI 2024 · 31 citations
- Ameliorate Spurious Correlations in Dataset CondensationJustin Cui, Ruochen Wang, Yuanhao Xiong, Cho-Jui HsiehICML 2024 · 7 citations
- Color-Oriented Redundancy Reduction in Dataset DistillationBowen Yuan, Zijian Wang, Mahsa Baktashmotlagh, Yadan Luo et al.NeurIPS 2024 · 7 citations
- Dataset Distillation as Data Compression: A Rate-Utility PerspectiveYouneng Bao, Yiping Liu, Zhuo Chen, Yongsheng Liang et al.ICCV 2025 · 3 citations
- FARTrack: Fast Autoregressive Visual Tracking with High PerformanceGuijie Wang, Tong Lin, Yifan Bai, Anjia Cao et al.ICLR 2026 · 3 citations
Builds on23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Image2StyleGAN: How to Embed Images Into the StyleGAN Latent Space?Rameen Abdal, Yipeng Qin, Peter WonkaICCV 2019 · 1,195 citations
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 684 citations
Related papers
- Frequency Domain-Based Dataset DistillationDongHyeok Shin, Seungjae Shin, Il-Chul MoonNeurIPS 2023 · 39 citations
- Distilling Dataset into Neural FieldDonghyeok Shin, HeeSun Bae, Gyuwon Sim, Wanmo Kang et al.ICLR 2025
- Understanding Dataset Distillation via Spectral FilteringDeyu Bo, Songhua Liu, Xinchao WangICLR 2026 · 3 citations
- Slimmable Dataset CondensationSonghua Liu, Jingwen Ye, Runpeng Yu, Xinchao WangCVPR 2023
- D4M: Dataset Distillation via Disentangled Diffusion ModelDuo Su, Junjie Hou, Weizhi Gao, Yingjie Tian et al.CVPR 2024 · 11 citations
