Sparse Parameterization for Epitomic Dataset Distillation
Xing Wei, Anjia Cao, Funing Yang, Zhiheng Ma
摘要
The success of deep learning relies heavily on large and diverse datasets, but the storage, preprocessing, and training of such data present significant challenges. To address these challenges, dataset distillation techniques have been proposed to obtain smaller synthetic datasets that capture the essential information of the originals. In this paper, we introduce a Sparse Parameterization for Epitomic datasEt Distilla-tion (SPEED) framework, which leverages the concept of dictionary learning and sparse coding to distill epitomes that represent pivotal information of the dataset. SPEED prioritizes proper parameterization of the synthetic dataset and introduces techniques to capture spatial redundancy within and between synthetic images. We propose Spatial-Agnostic Epitomic Tokens (SAETs) and Sparse Coding Matrices (SCMs) to efficiently represent and select significant features. Additionally, we build a Feature-Recurrent Network (FReeNet) to generate hierarchical features with high compression and storage efficiency. Experimental results demonstrate the superiority of SPEED in handling high-resolution datasets, achieving state-of-the-art performance on multiple benchmarks and downstream applications. Our framework is compatible with a variety of dataset matching approaches, generally enhancing their performance. This work highlights the importance of proper parameterization in epitomic dataset distillation and opens avenues for efficient representation learning. Source code is available at https://github.com/MIV-XJTU/SPEED .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Evolving Parameterized Prompt Memory for Continual LearningMuhammad Rifki Kurniawan, Xiang Song, Zhiheng Ma, Yuhang He 等AAAI 2024 · 被引用 31 次
- Ameliorate Spurious Correlations in Dataset CondensationJustin Cui, Ruochen Wang, Yuanhao Xiong, Cho-Jui HsiehICML 2024 · 被引用 7 次
- Color-Oriented Redundancy Reduction in Dataset DistillationBowen Yuan, Zijian Wang, Mahsa Baktashmotlagh, Yadan Luo 等NeurIPS 2024 · 被引用 7 次
- Dataset Distillation as Data Compression: A Rate-Utility PerspectiveYouneng Bao, Yiping Liu, Zhuo Chen, Yongsheng Liang 等ICCV 2025 · 被引用 3 次
- FARTrack: Fast Autoregressive Visual Tracking with High PerformanceGuijie Wang, Tong Lin, Yifan Bai, Anjia Cao 等ICLR 2026 · 被引用 3 次
它引用的顶会 Paper23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Image2StyleGAN: How to Embed Images Into the StyleGAN Latent Space?Rameen Abdal, Yipeng Qin, Peter WonkaICCV 2019 · 被引用 1,195 次
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 被引用 684 次
相关 Paper
- Frequency Domain-Based Dataset DistillationDongHyeok Shin, Seungjae Shin, Il-Chul MoonNeurIPS 2023 · 被引用 39 次
- Distilling Dataset into Neural FieldDonghyeok Shin, HeeSun Bae, Gyuwon Sim, Wanmo Kang 等ICLR 2025
- Understanding Dataset Distillation via Spectral FilteringDeyu Bo, Songhua Liu, Xinchao WangICLR 2026 · 被引用 3 次
- Slimmable Dataset CondensationSonghua Liu, Jingwen Ye, Runpeng Yu, Xinchao WangCVPR 2023
- D4M: Dataset Distillation via Disentangled Diffusion ModelDuo Su, Junjie Hou, Weizhi Gao, Yingjie Tian 等CVPR 2024 · 被引用 11 次
